Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Morphing into Hybrid Attention Models

About

Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer importance as isolated and overlooking the interdependent layer effect under a global hybrid configuration. In this work, we formulate hybrid layer selection as a budget-constrained subset optimization problem. We further propose FlashMorph (Fast LAyer Selection for Hybrid MORPHing), an effective, efficient and scalable layer selection method for Transformer-to-hybrid conversion. FlashMorph first constructs a morphable model by equipping each full-attention layer with a converted linear-attention branch. It then freezes all model weights and jointly optimizes layerwise gates on synthetic long-context retrieval data, with a linearization regularization that encourages the model to rely on linear attention for efficiency. The learned gates are discretized under a preset full-attention budget to instantiate the hybrid architecture, followed by standard logits distillation and long-context finetuning. Extensive experiments show that FlashMorph discovers more effective hybrid configurations, preserves strong long-context recall and general benchmark performance while substantially reducing layer selection cost compared with existing layer selection methods, demonstrating its effectiveness, efficiency, and scalability.

Disen Lan, Jianbin Zheng, Yuxi Ren, Xin Xia, Xuanda Wang, Xuefeng Xiao, Xipeng Qiu, Yu Cheng• 2026

Related benchmarks

TaskDatasetResultRank
Long-context recallLong-context Recall-intensive Tasks Suite SQuAD, FDA, SWDE
SQuAD Recall Accuracy54.3
48
Commonsense ReasoningCommonsense Reasoning Suite (PIQA, ARC-e, ARC-c, HellaSwag, WinoGrande)
PIQA Accuracy73.3
35
Commonsense ReasoningARC-e, ARC-c, PIQA, WinoG., HellaS.
ARC-e Accuracy81.9
25
Needle-In-A-Haystack RetrievalNIAH-Single-3 256K context
Accuracy94
23
Needle-In-A-Haystack RetrievalNIAH-Single-1 128K context
Accuracy100
23
Needle-In-A-Haystack RetrievalNIAH-Single 128K context 2
Accuracy98.2
23
Needle-In-A-Haystack RetrievalNIAH-Single-2 256K context
Accuracy88.2
23
Needle-In-A-Haystack RetrievalNIAH-Single-3 128K context
Accuracy94.4
23
Needle-In-A-Haystack RetrievalNIAH-Single-1 32K context
Accuracy (NIAH-Single-1 32K)100
23
Needle-In-A-Haystack RetrievalNIAH-Single-1 256K context
Accuracy100
23
Showing 10 of 17 rows

Other info

GitHub

Follow for update