Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

About

Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training.

Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, Caglar Gulcehre• 2024

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningWinoGrande
Accuracy65.2
1581
Commonsense ReasoningHellaSwag
HellaSwag Accuracy67.2
897
Physical Commonsense ReasoningPIQA
Accuracy77.4
724
Multitask Language UnderstandingMMLU
Accuracy29.5
568
Commonsense ReasoningPIQA
Accuracy81
400
Commonsense ReasoningARC Challenge--
259
Language ModelingPG-19--
244
Long-range sequence modelingLong Range Arena (LRA)
Text Accuracy71.75
177
Physical Commonsense ReasoningPIQA (val)
Accuracy66.1
118
Question AnsweringARC Challenge (test)
Accuracy25.4
103
Showing 10 of 36 rows

Other info

Follow for update