Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Abstraction and Reasoning on 1D-ARC
Loading...
397
Accuracy
IO-Prompt
-15.7656
91.3947
198.555
305.7153
Apr 14, 2026
Apr 26, 2026
May 8, 2026
May 20, 2026
Jun 1, 2026
Jun 13, 2026
Jun 25, 2026
Accuracy
Matching Accuracy
Match-and-Solve Accuracy
Updated 29d ago
Evaluation Results
Method
Method
Links
Accuracy
Matching Accuracy
Match-and-Solve Accuracy
IO-Prompt
Base model=gpt-4o-2024...
2026.04
397
-
-
Auto CoT-Prompt
Base model=gpt-4o-2024...
2026.04
346
-
-
Minitron + DIARC
#Params=6.9B
2026.06
58.82
-
-
Minitron + SFT only
#Params=6.9B
2026.06
57.05
-
-
Qwen-2.5
#Params=7B
2026.06
55
-
-
Llama-3.2 + DIARC
#Params=2.6B
2026.06
52.39
-
-
GPT-4
#Params=N/A
2026.06
52
-
-
Llama-3.2 + SFT only
#Params=2.6B
2026.06
51.72
-
-
Qwen3 + DIARC
#Params=3.6B
2026.06
45.95
-
-
Qwen3 + SFT only
#Params=3.6B
2026.06
42.84
-
-
Minitron
#Params=8B
2026.06
3.55
-
-
Llama-3.2
#Params=3B
2026.06
0.11
-
-
Qwen3
#Params=4B
2026.06
0.11
-
-
Feedback
Search any
task
Search any
task