| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Text-to-Image Generation | Qwen-Image | Image Reward1.295 | 96 | |
| Single-task capability preservation | Qwen-Image | DINO Score76.91 | 24 | |
| Text-to-Image Generation | Qwen-Image 1664 x 928 | Latency (s)6.95 | 13 | |
| Text-to-Image Generation | Qwen-Image Evaluation Set | Latency (s)22.4 | 12 | |
| Text-to-Image Generation | Qwen-Image native (test) | Speedup10.3 | 11 | |
| Image-level forgery detection | Qwen-Image | Accuracy82.1 | 9 | |
| Pixel-level image forgery localization | Qwen-Image (OOD) | IoU48.2 | 9 | |
| Identity Customization | Qwen-Image | Similarity0.703 | 8 | |
| AIGI Detection | Qwen-Image 12B 2.5 (test) | Accuracy99.07 | 7 | |
| AIGI Detection | Qwen-Image (test) | Accuracy97.05 | 7 | |
| Resolution extrapolation | Qwen-Image Direct extrapolation (test) | FID78.15 | 6 | |
| Style Customization | Qwen-Image | Similarity71 | 4 | |
| Text-to-Image Generation | Qwen-Image | FID28.13 | 3 | |
| Text-to-Image Generation | Qwen-Image-Lightning | Latency (s)10.91 | 3 | |
| Generative Diversity Evaluation | Qwen-Image | DINOv3 Score0.958 | 3 | |
| AI-generated image detection | Qwen-Image | Accuracy94.1 | 3 | |
| Human Evaluation | Qwen-Image rollout results | Win Rate77.5 | 2 | |
| Multimodal Inference | Qwen-Image (inference) | Inference Latency (s)14.92 | 2 |