Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Training Throughput on Synthetic 8192-token context (train)
Loading...
229,957
Training Throughput (tok/s)
Full-sequence attention
148,856.76
169,911.63
190,966.5
212,021.37
Jul 2, 2026
Training Throughput (tok/s)
Updated 18d ago
Evaluation Results
Method
Method
Links
Training Throughput (tok/s)
Full-sequence attention
GPUs=8x A100, Global b...
2026.07
229,957
Full-sequence DiffuMamba-H
GPUs=8x A100, Global b...
2026.07
223,825
Partially Reverse BDLM Mamba-H
GPUs=8x A100, Global b...
2026.07
166,849
BDLM attention
GPUs=8x A100, Global b...
2026.07
151,976
Feedback
Search any
task
Search any
task