Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Prefill Latency Measurement on LLaMA 8B-Instruct 64K context length 3.1
Loading...
8,945.9
Attention Latency
CompactAttention-FP
8,310.968
12,596.759
16,882.55
21,168.341
May 16, 2026
Attention Latency
End-to-End Latency
Updated 2mo ago
Evaluation Results
Method
Method
Links
Attention Latency
End-to-End Latency
CompactAttention-FP
GPU=RTX PRO 6000, TP=2...
2026.05
8,945.9
23,707.8
CompactAttention-SA
GPU=RTX PRO 6000, TP=2...
2026.05
13,823.5
28,960.4
FlashPrefill
GPU=RTX PRO 6000, TP=2...
2026.05
14,381.8
29,411.7
QUOKA
GPU=RTX PRO 6000, TP=2...
2026.05
16,325.4
30,588
SeerAttention
GPU=RTX PRO 6000, TP=2...
2026.05
17,078.2
32,545.8
Dense
GPU=RTX PRO 6000, TP=2...
2026.05
17,838.8
32,177.4
XAttention
GPU=RTX PRO 6000, TP=2...
2026.05
24,819.2
39,206.6
Feedback
Search any
task
Search any
task