Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Reasoning failure prediction on CodeLingua (L2)

75Accuracy

thought-tree-based classifier

67.7269.6171.573.39Apr 18, 2026
Updated 1mo ago

Evaluation Results

MethodLinks
2026.04
75
2026.04
68