Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
About
We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully design the transformer structure, text encoder, and positional encoding. We also build from scratch a whole data pipeline to update and evaluate data for iterative model optimization. For fine-grained language understanding, we train a Multimodal Large Language Model to refine the captions of the images. Finally, Hunyuan-DiT can perform multi-turn multimodal dialogue with users, generating and refining images according to the context. Through our holistic human evaluation protocol with more than 50 professional human evaluators, Hunyuan-DiT sets a new state-of-the-art in Chinese-to-image generation compared with other open-source models. Code and pretrained models are publicly available at github.com/Tencent/HunyuanDiT
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text-to-Image Generation | GenEval | Overall Score63 | 467 | |
| Text-to-Image Generation | GenEval | GenEval Score63 | 277 | |
| Text-to-Image Generation | DPG-Bench | Overall Score78.87 | 173 | |
| Text-to-Image Generation | DPG | Overall Score78.87 | 131 | |
| Text-to-Image Generation | DPG-Bench (test) | Global Fidelity84.59 | 43 | |
| Text-to-Image Generation | GenAI-Bench | Average Score0.721 | 30 | |
| Text-to-Image Generation | HPS v3 | Overall Score8.19 | 24 | |
| Text-to-Image Alignment | DPG | Entity80.59 | 21 | |
| Text-to-Image Generation | DrawBench | VQAScore0.712 | 18 | |
| Text-to-Image Generation | TIIF Bench mini (test) | Overall Score (Short)51.38 | 18 |