Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Decode Throughput on Gemma 26B-A4B 4
Loading...
69.3
Throughput (tok/s)
MLX
57.548
60.599
63.65
66.701
Jul 1, 2026
Throughput (tok/s)
Speedup vs. llama.cpp
Ratio vs. MLX
Updated 24d ago
Evaluation Results
Method
Method
Links
Throughput (tok/s)
Speedup vs. llama.cpp
Ratio vs. MLX
MLX
Quantization=Q4, Hardw...
2026.07
69.3
-
-
BaseRT
Quantization=Q4, Hardw...
2026.07
62.2
1.07
0.9
llama.cpp
Quantization=Q4, Hardw...
2026.07
58
-
-
Feedback
Search any
task
Search any
task