Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Language Model Inference on vLLM benchmark (1024 prompts)
Loading...
1.839
Latency (s)
Fisher-IntDim-E
1.82324
1.92962
2.036
2.14238
Jun 4, 2026
Latency (s)
Throughput (tok/s)
KV-cache Budget (K)
VRAM (GiB)
Active Params (B)
Updated 1mo ago
Evaluation Results
Method
Method
Links
Latency (s)
Throughput (tok/s)
KV-cache Budget (K)
VRAM (GiB)
Active Params (B)
Fisher-IntDim-E
Hardware=NVIDIA H100,...
2026.06
1.839
8,909
251
15.09
2.274
Qwen1.5-MoE
Hardware=NVIDIA H100,...
2026.06
2.233
7,336
187
26.67
2.689
Feedback
Search any
task
Search any
task