Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Core Context Aware Transformers for Long Context Language Modeling

About

Transformer-based Large Language Models (LLMs) have exhibited remarkable success in extensive tasks primarily attributed to self-attention mechanism, which requires a token to consider all preceding tokens as its context to compute attention. However, when the context length L becomes very large (e.g., 128K), the amount of potentially redundant information in the context tends to increase. The redundant context not only hampers the modeling representation performance but also incurs unnecessary computational and storage overhead. In this paper, we propose a plug-and-play Core Context Aware (CCA) Attention for efficient long-context modeling, comprising two complementary modules: 1) Globality-aware pooling module groups input tokens and dynamically compresses each group into one core token based on their significance. In this way, our method automatically focuses and strengthens core context while diminishing redundancy during the learning process, leading to effective long-term dependency modeling. 2) Locality-preserving module incorporates neighboring tokens to preserve local context for detailed representation. Notably, our CCA-Attention is able to replace the self-attention module in existing LLMs with minimal fine-tuning cost. Extensive experimental results show the superiority of our method in both long-context modeling and computational efficiency over state-of-the-art methods.

Yaofo Chen, Zeng You, Shuhai Zhang, Haokun Li, Yirui Li, Yaowei Wang, Mingkui Tan• 2024

Related benchmarks

TaskDatasetResultRank
Long-context Language UnderstandingRULER 32k context length
FWE12.83
39
Long-context Language UnderstandingRULER 16k context length
FWE Score29.83
21
Long-context Language UnderstandingRULER 4k context length
FWE Rate45.67
16
Long-context UnderstandingRULER 8k context
CWE36.25
13
Long-context Language UnderstandingLongBench-E 2024 (test)
Short Context QA Score6.22
12
Long-context Information ExtractionRULER 4K-32K Average
CWE Score24.9
6
Showing 6 of 6 rows

Other info

Follow for update