Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

AMA-Bench

Benchmarks

Task NameDataset NameSOTA ResultTrend
State AbstractionAMA-Bench SA
Success Metric52.92
20
State UpdatingAMA-Bench (SU)
Success Metric58.73
20
Causal InferenceAMA-Bench (CI)
Success Metric57.21
20
Agent MemoryAMA-Bench real-world
Recall (Accuracy)62.38
14
Embodied AIAMA-bench Embodied AI Domain
TTFT Mean (std) [s]0.137
8
Trajectory QAAMA-Bench Full 208-episode trajectory
F1 Score38.2
5
Showing 6 of 6 rows