Share your thoughts, 1 month free Claude Pro on usSee more

Language Modeling Inference on Qwen2.5-7B (256K Context Length)

26.3Decode Latency (ms/token)

FastMKA

Updated 4mo ago

Evaluation Results

Method	Links
FastMKA 2026.03		26.3	1.86
MLA 2026.03		48.9	-
GQA 2026.03		75.2	-
MHA 2026.03		87.6	-