Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

About

We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully simulate all future pixels, but should learn and maintain the information most relevant to embodiment control: object state, spatial relations, contact conditions, task progress, action consequences, failure boundaries, and deployment uncertainty. Kairos establishes three model-side prerequisites toward this goal. First, it \textbf{learns} control-relevant information through a \textbf{Cross-Embodiment Data Curriculum}, which organizes open-world videos, human behavioral data, and robot interactions into an intervention-strength progression from passive physical observation to intentional behavior and embodied action grounding. Second, it \textbf{maintains} control-sufficient states through a unified \textbf{understanding, generation, and prediction architecture} equipped with \textbf{Hybrid Linear Temporal Attention}, where local, mid-range, and global temporal pathways support multi-timescale state maintenance under efficient inference. Third, it \textbf{deploys} these states through a \textbf{Deployment-Aware System Co-Design}, treating latency, memory footprint, and hardware compatibility as first-order constraints for future observation, action, and feedback loops. Experiments on embodied world-model benchmarks, world-action benchmarks, long-horizon generation, and inference-efficiency evaluation show that Kairos achieves superior performance while offering a favorable efficiency to capability trade-off.

Kairos Team: Fei Wang, Shan You, Qiming Zhang, Tao Huang, Zuoyi Fu, Zhisheng Zheng, Yunlong Xi, Feng Lv, Xiaoming Wu, Zeyu Liu, Cong Wan, Pu Li, Ruiqing Yang, Xiaoou Li, Wei Wang, Kangkang Zhu, Yuwei Zhang, Shi Fu, Zheng Zhang, Xiaoning Wu, Xuzeng Fan, Dacheng Tao, Xiaogang Wang• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationLIBERO-Plus
Language Understanding Score95.3
414
Video GenerationVideoPhy--
50
Bimanual ManipulationRoboTwin Clean setting 2.0
Success Rate96.9
36
Embodied Video GenerationPAI-Bench robot domain
Domain Score88.59
36
Bimanual ManipulationRoboTwin 2.0 (random)
Success Rate95.2
26
Bimanual ManipulationRoboTwin 2.0
Success Rate96.1
25
Robotics Video GenerationDreamGen Bench
GR1 Object Score (Qwen-IF)66
15
Video GenerationWorldModelBench
Instruction Score2.36
11
World ModelingWorldModelBench Robot Set
Instruction Following Score2.36
8
Physical AI generationPAI-Bench robotics subset
I2V Background Score97.75
4
Showing 10 of 13 rows

Other info

GitHub

Follow for update