Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

About

Recent Large Reasoning Models significantly improve the reasoning ability of Large Language Models by learning to reason, exhibiting the promising performance in solving complex tasks. LRMs solve tasks that require complex reasoning by explicitly generating reasoning trajectories together with answers. Nevertheless, judging the quality of such an output answer is not easy because only considering the correctness of the answer is not enough and the soundness of the reasoning trajectory part matters as well. Logically, if the soundness of the reasoning part is poor, even if the answer is correct, the confidence of the derived answer should be low. Existing methods did consider jointly assessing the overall output answer by taking into account the reasoning part, however, their capability is still not satisfactory as the causal relationship of the reasoning to the concluded answer cannot properly reflected. In this paper, inspired by classical mechanics, we present a novel approach towards establishing a CoT-Kinetics energy equation. Specifically, our CoT-Kinetics energy equation formulates the token state transformation process, which is regulated by LRM internal transformer layers, as like a particle kinetics dynamics governed in a mechanical field. Our CoT-Kinetics energy assigns a scalar score to evaluate specifically the soundness of the reasoning phase, telling how confident the derived answer could be given the evaluated reasoning. As such, the LRM's overall output quality can be accurately measured, rather than a coarse judgment (e.g., correct or incorrect) anymore.

Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma• 2025

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH 500
Accuracy82.6
589
Mathematical ReasoningAIME 2024
Accuracy26.7
525
Mathematical ReasoningAMC
Accuracy (%)64.8
375
Mathematical ReasoningAIME 2025
Accuracy22.7
353
ReasoningMMLU-Pro
Accuracy28.9
264
Mathematical ReasoningMinerva Math
Accuracy29.7
124
Mathematical ReasoningIn-Distribution Reasoning Performance Suite (AIME, AMC, MATH-500, Minerva, Olympiad)
AIME 2024 Score31.9
119
Mathematical ReasoningOlympiadBench Math
Accuracy45.7
97
Uncertainty EstimationJudgeBench (test)
AUROC63.72
77
General ReasoningOut-of-Distribution Performance Suite (ARC-c, GPQA*, MMLU-Pro) (test)
ARC-c Score15.6
73
Showing 10 of 21 rows

Other info

Follow for update