Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Prefill Latency Measurement on LLaMA 8B Instruct 128K context length 3.1
Loading...
24,055.4
Attention Latency (ms)
CompactAttention-FP
21,712.936
37,524.568
53,336.2
69,147.832
May 16, 2026
Attention Latency (ms)
End-to-End Latency (ms)
Updated 2mo ago
Evaluation Results
Method
Method
Links
Attention Latency (ms)
End-to-End Latency (ms)
CompactAttention-FP
GPU=RTX PRO 6000, TP=2...
2026.05
24,055.4
54,130.6
CompactAttention-SA
GPU=RTX PRO 6000, TP=2...
2026.05
32,846
64,126.7
FlashPrefill
GPU=RTX PRO 6000, TP=2...
2026.05
43,750.5
80,108.5
SeerAttention
GPU=RTX PRO 6000, TP=2...
2026.05
49,839.8
81,774.4
QUOKA
GPU=RTX PRO 6000, TP=2...
2026.05
56,475.3
85,375.3
Dense
GPU=RTX PRO 6000, TP=2...
2026.05
65,971.1
94,944.3
XAttention
GPU=RTX PRO 6000, TP=2...
2026.05
82,617
113,016.5
Feedback
Search any
task
Search any
task