Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Throughput Benchmarking on vLLM (1k Input / 1k Output Tokens)
Loading...
2.42
Throughput (req/s)
Fisher-MoE
1.0992
1.4421
1.785
2.1279
Jun 4, 2026
Throughput (req/s)
Throughput (tok/s)
Speedup vs Qwen1.5-7B
Speedup vs Qwen1.5-MoE-A2.7B
Updated 1mo ago
Evaluation Results
Method
Method
Links
Throughput (req/s)
Throughput (tok/s)
Speedup vs Qwen1.5-7B
Speedup vs Qwen1.5-MoE-A2.7B
Fisher-MoE
Hardware=NVIDIA A100-8...
2026.06
2.42
4,853.52
2.1
1.21
Qwen1.5-MoE-A2.7B-Chat
Hardware=NVIDIA A100-8...
2026.06
2.01
4,010.27
1.74
1
Qwen1.5-7B-Chat
Hardware=NVIDIA A100-8...
2026.06
1.15
2,298.89
1
0.57
Feedback
Search any
task
Search any
task