Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

About

The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key-value (KV) cache, quickly overwhelming GPU memory. To address this challenge, we propose ThinKV, a thought-adaptive KV cache compression framework. ThinKV is based on the observation that attention sparsity reveals distinct thought types with varying importance within the CoT. It applies a hybrid quantization-eviction strategy, assigning token precision by thought importance and progressively evicting tokens from less critical thoughts as reasoning trajectories evolve. Furthermore, to implement ThinKV, we design a kernel that extends PagedAttention to enable efficient reuse of evicted tokens' memory slots, eliminating compaction overheads. Extensive experiments on DeepSeek-R1-Distill, GPT-OSS, and NVIDIA AceReason across mathematics and coding benchmarks show that ThinKV achieves near-lossless accuracy with less than 5% of the original KV cache, while improving performance with up to 5.8x higher inference throughput over state-of-the-art baselines.

Akshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan, Brucek Khailany, Tushar Krishna• 2025

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH 500
Accuracy71
589
Mathematical ReasoningGSM8K
Accuracy (Acc)60.1
352
Mathematical ReasoningAIME 2024 (test)
Accuracy30
294
ReasoningLiveCodeBench
pass@1 Accuracy50.47
71
Math ReasoningAIME 2024
Mean Wall-Clock Time (s)721
61
Mathematical ReasoningMATH 500
Mean Wall-Clock Time344.6
49
Math problem solvingAIME 2024
Mean Peak GPU Memory (MB)1.55e+4
41
Mathematical ReasoningMATH 500
Mean Peak GPU Memory (MB)1.55e+4
37
Mathematical ReasoningAIME
Accuracy33.2
18
Mathematical ReasoningMATH 500
Accuracy76
18
Showing 10 of 15 rows

Other info

Follow for update