Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Stabilizing Native Low-Rank LLM Pretraining

About

Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges. Low-rank factorization offers a promising route to reduce training and inference costs, but the community lacks a stable recipe for training models from scratch using exclusively low-rank weights while matching the performance of the dense model. We demonstrate that Large Language Models (LLMs) can be trained from scratch using exclusively low-rank factorized weights for all non-embedding matrices without auxiliary "full-rank" guidance required by prior methods. While native low-rank training often suffers from instability and loss spikes, we identify uncontrolled growth in the spectral norm (largest singular value) of the weight matrix update as the dominant factor. To address this, we introduce Spectron: Spectral renormalization with orthogonalization, which dynamically bounds the resultant weight updates based on the current spectral norms of the factors. Our method enables stable, end-to-end factorized training with negligible overhead. Finally, we establish compute-optimal scaling laws for natively low-rank transformers, demonstrating predictable power-law behavior and improved inference efficiency relative to dense models.

Paul Janson, Edouard Oyallon, Eugene Belilovsky• 2026

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningHellaSwag
Accuracy40.11
1896
Physical Commonsense ReasoningPIQA
Accuracy66.76
696
Question AnsweringARC Easy
Accuracy36.78
597
Natural Language UnderstandingGLUE (val)
SST-297.25
201
Natural language generationE2E (test)
ROUGE-L71.68
100
Table-to-text generationE2ENLG (test)
BLEU70.26
51
Matrix completionMovieLens 1M (test)
RMSE0.8691
37
Language ModelingFineWeb 100M token (val)
Perplexity12.11
9
Natural Language UnderstandingGLUE (test)
MNLI Score89.77
5
Showing 9 of 9 rows

Other info

Follow for update