Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models

About

State Space Models (SSMs) such as Mamba-2 offer linear-time inference but their memory footprint limits edge deployment. Prior ternary SSM work (Slender-Mamba) trains from scratch on 150B tokens; we show a pretrained checkpoint suffices, reducing the marginal token budget by 1,000x. Using grouped quantization-aware training (QAT) with knowledge distillation from a frozen FP16 teacher, we compress Mamba-2 1.3B to 3.61x (2,687 to 744 MB) and achieve 48.1% zero-shot accuracy (7-task average) in just 102M tokens (4 GPU-hours, single H100) -- approaching Bi-Mamba's 48.4% (within +/-0.9pp CI). This QAT-from-pretrained setting reveals zero-ratio collapse, a novel instability caused by learnable quantization scales that does not arise in from-scratch training. We further show that post-hoc correction strategies effective for Transformers fail for SSMs due to error accumulation through the recurrence. These results demonstrate that ternary SSMs do not require expensive from-scratch training: QAT from pretrained checkpoints with KD is a data-efficient alternative.

Ramprasath Ganesaraja, Sahil Dilip Panse, Swathika N• 2026

Related benchmarks

TaskDatasetResultRank
Language ModelingWikiText-2 (test)
PPL19.63
2416
Language ModelingC4
Perplexity20.99
482
Zero-shot language evaluation7-task evaluation suite (BoolQ, PIQA, HellaSwag, WinoGrande, ARC-E, ARC-C, OBQA) (test)
BoolQ Accuracy56
4
Inference Memory AnalysisNVIDIA H100 NVL seq_len=1024
Model Weights Size744
3
Zero-shot Evaluation7-task suite
Average Score (7-task Suite)48.1
3
Zero-shot Classification7-task suite
Average Accuracy (7-task)48.1
2
Showing 6 of 6 rows

Other info

Follow for update