Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Question Answering on MMLU (Machine Learning subset, test split)
Loading...
88.4
Accuracy
TEXTGRAD
52.5304
61.8427
71.155
80.4673
Jun 11, 2024
Oct 13, 2024
Feb 14, 2025
Jun 19, 2025
Oct 21, 2025
Feb 22, 2026
Jun 27, 2026
Accuracy
Updated 25d ago
Evaluation Results
Method
Method
Links
Accuracy
TEXTGRAD
Mode=Zero-shot, Model=...
2024.06
88.4
CoT
Mode=Zero-shot, Model=...
2024.06
85.7
LLM-based QA agent with external memory
Model Backend=GPT 5.4...
2026.06
85.16
LLM-based QA agent with external memory
Model Backend=GPT-4o M...
2026.06
70.31
LLM-based QA agent with external memory
Model Backend=Gemma2-9...
2026.06
56.25
LLM-based QA agent with external memory
Model Backend=Phi3-14B...
2026.06
53.91
Feedback
Search any
task
Search any
task