Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Decode Throughput on Qwen3 0.6B
Loading...
464.5
Throughput (tok/s)
BaseRT
210.012
276.081
342.15
408.219
Jul 1, 2026
Throughput (tok/s)
Speedup vs. llama.cpp
Speedup vs. MLX
Updated 24d ago
Evaluation Results
Method
Method
Links
Throughput (tok/s)
Speedup vs. llama.cpp
Speedup vs. MLX
BaseRT
Quantization=Q4, Hardw...
2026.07
464.5
1.56
1.35
MLX
Quantization=Q4, Hardw...
2026.07
343.6
-
-
BaseRT
Quantization=Q8, Hardw...
2026.07
321.2
1.46
1.26
llama.cpp
Quantization=Q4, Hardw...
2026.07
297.4
-
-
MLX
Quantization=Q8, Hardw...
2026.07
255.3
-
-
llama.cpp
Quantization=Q8, Hardw...
2026.07
219.8
-
-
Feedback
Search any
task
Search any
task