Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
About
We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully simulate all future pixels, but should learn and maintain the information most relevant to embodiment control: object state, spatial relations, contact conditions, task progress, action consequences, failure boundaries, and deployment uncertainty. Kairos establishes three model-side prerequisites toward this goal. First, it \textbf{learns} control-relevant information through a \textbf{Cross-Embodiment Data Curriculum}, which organizes open-world videos, human behavioral data, and robot interactions into an intervention-strength progression from passive physical observation to intentional behavior and embodied action grounding. Second, it \textbf{maintains} control-sufficient states through a unified \textbf{understanding, generation, and prediction architecture} equipped with \textbf{Hybrid Linear Temporal Attention}, where local, mid-range, and global temporal pathways support multi-timescale state maintenance under efficient inference. Third, it \textbf{deploys} these states through a \textbf{Deployment-Aware System Co-Design}, treating latency, memory footprint, and hardware compatibility as first-order constraints for future observation, action, and feedback loops. Experiments on embodied world-model benchmarks, world-action benchmarks, long-horizon generation, and inference-efficiency evaluation show that Kairos achieves superior performance while offering a favorable efficiency to capability trade-off.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robotic Manipulation | LIBERO-Plus | Language Understanding Score95.3 | 414 | |
| Video Generation | VideoPhy | -- | 50 | |
| Bimanual Manipulation | RoboTwin Clean setting 2.0 | Success Rate96.9 | 36 | |
| Embodied Video Generation | PAI-Bench robot domain | Domain Score88.59 | 36 | |
| Bimanual Manipulation | RoboTwin 2.0 (random) | Success Rate95.2 | 26 | |
| Bimanual Manipulation | RoboTwin 2.0 | Success Rate96.1 | 25 | |
| Robotics Video Generation | DreamGen Bench | GR1 Object Score (Qwen-IF)66 | 15 | |
| Video Generation | WorldModelBench | Instruction Score2.36 | 11 | |
| World Modeling | WorldModelBench Robot Set | Instruction Following Score2.36 | 8 | |
| Physical AI generation | PAI-Bench robotics subset | I2V Background Score97.75 | 4 |