OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models

About

The advent of large language models (LLMs) has revolutionized natural language processing, enabling unprecedented capabilities in understanding and generating human-like text. However, the computational cost and convergence times associated with fine-tuning these models remain significant challenges. Low-Rank Adaptation (LoRA) has emerged as a promising method to mitigate these issues by introducing efficient fine-tuning techniques with a reduced number of trainable parameters. In this paper, we present OLoRA, an enhancement to the LoRA method that leverages orthonormal matrix initialization through QR decomposition. OLoRA significantly accelerates the convergence of LLM training while preserving the efficiency benefits of LoRA, such as the number of trainable parameters and GPU memory footprint. Our empirical evaluations demonstrate that OLoRA not only converges faster but also exhibits improved performance compared to standard LoRA across a variety of language modeling tasks. This advancement opens new avenues for more efficient and accessible fine-tuning of LLMs, potentially enabling broader adoption and innovation in natural language applications.

Kerim B\"uy\"ukaky\"uz• 2024

Related benchmarks

Task	Dataset	Result
Math Reasoning	GSM8K	Accuracy52.38	254
Commonsense Reasoning	Commonsense Reasoning (BoolQ, PIQA, SIQA, HellaS., WinoG., ARC-e, ARC-c, OBQA)	BoolQ Accuracy71.34	245
Math Reasoning	MATH	Accuracy8.22	160
Multi-turn dialogue	MT-Bench	MT-Bench Score6.13	126
Chat	MT-Bench	MT-Bench Score4.99	73
Sentiment Analysis	CR	CA93.16	54
Commonsense Reasoning	Commonsense Reasoning Benchmark	BoolQ Accuracy74.41	22
Language Modeling	Wikitext-2 raw v1	Loss2.241	10
Natural Language Understanding	GLUE	MRPC Score87.99	10
Mathematical Reasoning	GSM8K (test)	Loss0.503	10

Showing 10 of 16 rows

Other info

Follow for update

@wizwand_team Discord