Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM Inference Performance on LLM Inference Service Benchmark
Loading...
256
Max Concurrency
YouZhi-7B
46.96
101.23
155.5
209.77
Jun 4, 2026
Max Concurrency
Max Throughput (tokens/s)
Updated 1mo ago
Evaluation Results
Method
Method
Links
Max Concurrency
Max Throughput (tokens/s)
YouZhi-7B
KV Cache Size=576 elem...
2026.06
256
5,865
YiZhao-12B-Chat
KV Cache Size=512 elem...
2026.06
206
3,692
DianJin-R1-7B
KV Cache Size=1024 ele...
2026.06
190
5,445
YouZhi-14B
KV Cache Size=576 elem...
2026.06
134
2,990
Qwen3-8B
KV Cache Size=2048 ele...
2026.06
115
3,442
Qwen3.5-9B
KV Cache Size=2048 ele...
2026.06
102
3,040
OpenPangu-7B
KV Cache Size=2048 ele...
2026.06
95
3,325
Qwen2.5-14B-Ins.
KV Cache Size=2048 ele...
2026.06
55
1,740
Feedback
Search any
task
Search any
task