Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Decode Throughput on Gemma 4 E2B
Loading...
127.7
Throughput (tok/s)
BaseRT
56.772
75.186
93.6
112.014
Jul 1, 2026
Throughput (tok/s)
Speedup Ratio vs. llama.cpp
Speedup Ratio vs. MLX
Updated 24d ago
Evaluation Results
Method
Method
Links
Throughput (tok/s)
Speedup Ratio vs. llama.cpp
Speedup Ratio vs. MLX
BaseRT
Quantization=Q4, Hardw...
2026.07
127.7
1.19
-
llama.cpp
Quantization=Q4, Hardw...
2026.07
107
-
-
BaseRT
Quantization=Q8, Hardw...
2026.07
84.5
1.42
-
llama.cpp
Quantization=Q8, Hardw...
2026.07
59.5
-
-
Feedback
Search any
task
Search any
task