Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LLM-SRBENCH

Benchmarks

Task NameDataset NameSOTA ResultTrend
Symbolic RegressionLLM-SRBench (phys_osc)
Best Reward8.9985
16
Symbolic RegressionLLM-SRBench matsci
Best Reward8.25
16
Symbolic RegressionLLM-SRBench chem_react
Best Reward9
16
Symbolic RegressionLLM-SRBench bio_pop_growth
Best Reward8.9725
16
Symbolic RegressionLLM-SRBench Symbolic
Term Recall34.4
14
Symbolic RegressionLLM-SRBench OOD (test)
NMSE0.325
14
Symbolic RegressionLLM-SRBench ID (test)
NMSE0.4
14
Symbolic RegressionLLM-SRBENCH LSR-Transform
NMSE0.067
13
Numerical Symbolic RegressionLLM-SRBench 129-task synthetic (test)
Chemistry 95% Acc (Tol=0.01)35
11
Symbolic RegressionLLM-SRBench LSR-Synth Biology
NMSE0.64
10
Symbolic RegressionLLM-SRBench LSR-Synth Chemistry
NMSE0.0002
10
Symbolic RegressionLLM-SRBENCH LSR-Syn
Chemistry Error0
9
Symbolic RegressionLLM-SRBench (official 239-problem split)
Acc0.1 (%)77
6
Symbolic RegressionLLM-SRBench 129-task synthetic subset OOD (test)
Chemistry Accuracy (Tol 0.01)32
5
Symbolic RegressionLLM-SRBench Overall
Median R20.875
3
Symbolic RegressionLLM-SRBench Transform
Median R21.253
3
Symbolic RegressionLLM-SRBench Synthetic
Median R20.984
3
Showing 17 of 17 rows