Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Large Language Model Inference on Prompt length-1
Loading...
0.05
TTFT
CPU TEE
-9.108
52.7085
114.525
176.3415
Jun 16, 2026
TTFT
Total Latency
Decode Latency
Updated 1mo ago
Evaluation Results
Method
Method
Links
TTFT
Total Latency
Decode Latency
CPU TEE
Model Architecture=GPT...
2026.06
0.05
0.33
0.02
CPU TEE
Model Architecture=Qwe...
2026.06
0.11
0.52
0.06
Bifrost+
Model Architecture=GPT...
2026.06
3.9
62
3.9
Bifrost
Model Architecture=GPT...
2026.06
7.8
66
3.9
Bifrost+
Model Architecture=Qwe...
2026.06
22
177
22
Bifrost
Model Architecture=Qwe...
2026.06
45
202
22
Pure FHE
Model Architecture=GPT...
2026.06
68
579
34
Pure FHE
Model Architecture=Qwe...
2026.06
229
1,031
115
Feedback
Search any
task
Search any
task