| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Embodied Video Generation | PAI-Bench robot domain | Domain Score93.26 | 36 | |
| Image2World generation | PAI-Bench | Domain Score Average0.86 | 17 | |
| Text2World (T2W) Generation | PAI-Bench | Latency (s)26.28 | 14 | |
| Video spatial & 4D reasoning | PAI-Bench 2025 | Accuracy73.3 | 12 | |
| Physical Perception | PAI-Bench | PAI-Bench Score68.5 | 9 | |
| Image-to-Video | PAI-Bench VBench (test) | Delta Vote (%)0 | 8 | |
| Text-to-Video | PAI-Bench VBench (test) | Delta Vote (%)0 | 8 | |
| Image-to-Video | PAI-Bench VideoAlign (test) | ∆-Vote (%)0 | 8 | |
| Text-to-Video | PAI-Bench VideoAlign (test) | Delta Vote (%)0 | 8 | |
| Robotics Image-to-Video Generation | PAI-Bench-G | Grasp Success89.6 | 8 | |
| Physical AI generation | PAI-Bench robotics subset | I2V Background Score97.75 | 4 | |
| Video Generation | PAI-bench 15s | I2V Background Score97.9 | 4 | |
| Video Generation | PAI-Bench (full set) | I2V Background Fidelity97.9 | 4 | |
| Video spatial & 4D reasoning | PAI-Bench | Accuracy68.1 | 3 | |
| Text-to-World | PAI-Bench | Domain Average78.62 | 3 |