Context length extension¶
Goal¶
Increase the model's usable context window.
Examples:
- 4K → 8K
- 8K → 32K
- 32K → 128K
- 128K → 1M in specialized systems
Llama 3's largest model is reported with context up to 128K tokens. Gemini 1.5 and several research models reach 1M+. The frontier keeps moving.
Why long context is hard¶
Attention is expensive. Vanilla self-attention is quadratic in sequence length. Doubling context can roughly quadruple attention cost.
FlashAttention (Dao et al., 2022) helps by computing exact attention with fewer high-bandwidth memory reads/writes, improving speed and memory efficiency. It does not change asymptotic complexity but slashes the constant factor and enables much longer sequences in practice.
KV-cache memory at inference is the second bottleneck. At 128K context, KV cache for a 70B-class model can be tens of gigabytes per request — see Compression for mitigations (paged attention, quantized KV, GQA).
Main approaches¶
1. Train from scratch at long context¶
Best quality, very expensive.
Pros:
- Native long-context ability
- Fewer extrapolation issues
Cons:
- High compute cost
- Harder batching
- More memory pressure
Few labs do this end-to-end; most bootstrap from a shorter-context base model and extend.
2. RoPE scaling / interpolation¶
For RoPE-based models, modify positional frequency scaling so the model can operate beyond its original context.
Methods include:
- Linear (Position) Interpolation — Chen et al., 2023. Compresses positions linearly into the original range. Simple, works decently. arxiv.org/abs/2306.15595
- NTK-aware scaling — adjusts the RoPE base θ to interpolate high-frequency dimensions less. Works well with no fine-tuning for moderate extensions.
- YaRN — Peng et al., 2023. Dimension-dependent scaling. State-of-the-art with substantially fewer training tokens. arxiv.org/abs/2309.00071
- LongRoPE — Microsoft, 2024. Pushes to 2M+ context. arxiv.org/abs/2402.13753
YaRN specifically targets efficient context-window extension for RoPE models and reports requiring substantially fewer tokens and training steps than previous methods.
Default starting point
Pick YaRN unless you have a specific reason not to. It is the most-cited, best-supported method in 2025 open-weights work, and major inference frameworks (vLLM, SGLang, llama.cpp) implement it natively.
3. Continued pretraining on long sequences¶
After changing the positional scheme, train on long-context data.
Data examples:
- Books
- Long documents
- Legal contracts
- Code repositories
- Multi-file code tasks
- Long conversations
- Retrieval-augmented synthetic tasks
- Needle-in-a-haystack examples
- Long-context QA
Token budget for the long-context phase is typically small — billions of tokens, not trillions — but with careful mixing.
4. Architectural changes¶
Examples:
- Sliding-window attention — Mistral 7B v0.1 used 4K window with rolling buffer.
- Global tokens — a few special tokens that attend everywhere.
- Sparse attention — block-sparse, longformer-style.
- Memory tokens — learned compressed representations.
- Recurrence — Transformer-XL, RWKV-style.
- State-space hybrids — Mamba, Jamba, RecurrentGemma.
- Retrieval-augmented generation — see RAG.
- External memory — explicit retrieval over past content.
Long-context training data¶
You need tasks that require using far-away information.
Bad long-context data:
- Just concatenated random documents
- Long text where the answer is local
- Repetitive filler
- Synthetic "needle" tasks only
Good long-context data:
- Multi-section document QA
- Cross-document synthesis
- Long codebase navigation
- Legal/clinical/scientific evidence aggregation
- Conversation memory tasks
- Temporal consistency tasks
- Multi-hop retrieval within context
Needle-only training is a trap
A model trained only on synthetic needle-in-a-haystack examples can ace needle benchmarks while still failing real long-context tasks. The reasoning skills that matter (multi-hop synthesis, cross-section consistency, "the answer is implied across three places") need explicit training data.
Long-context evals¶
Evaluate:
- Needle-in-a-haystack
- Multi-needle retrieval
- LongBench-style tasks
- Book QA
- Long summarization
- Repository-level code understanding
- Long conversation consistency
- Lost-in-the-middle behavior
Important: passing needle retrieval does not mean the model can reason well over long context.
Eval suites worth running
- RULER — multi-dimensional long-context benchmark. arxiv.org/abs/2404.06654
- LongBench — multi-task Chinese/English long-context suite. arxiv.org/abs/2308.14508
- InfiniteBench — 100K+ token evals. arxiv.org/abs/2402.13718
- Needle-in-a-Haystack — Greg Kamradt's classic, still useful as a sanity check. github.com/gkamradt/LLMTest_NeedleInAHaystack
Practical tips¶
- Don't trust nominal context length. Many models claim 128K but degrade noticeably past 32K. Run RULER or LongBench, not just needle.
- Test "lost in the middle". Models often answer well from the start and end of the context but miss information from the middle. See Liu et al., 2023.
- Tune the inference stack. Long context exposes serving bugs (KV cache eviction, attention numerical issues, prefix caching). Test the served model, not just training-time behavior.
- Mix short and long during training. Pure long-context fine-tuning regresses short-context behavior; a mix of sequence lengths is essential.
Further reading¶
- Dao et al., "FlashAttention", 2022. arxiv.org/abs/2205.14135
- Dao, "FlashAttention-2", 2023. arxiv.org/abs/2307.08691
- Chen et al., "Extending Context Window of LLMs via Positional Interpolation", 2023. arxiv.org/abs/2306.15595
- Peng et al., "YaRN: Efficient Context Window Extension", 2023. arxiv.org/abs/2309.00071
- Ding et al., "LongRoPE", 2024. arxiv.org/abs/2402.13753
- Liu et al., "Lost in the Middle", 2023. arxiv.org/abs/2307.03172
- Hsieh et al., "RULER", 2024. arxiv.org/abs/2404.06654
- Llama 3 paper, §5 (long-context training), 2024. arxiv.org/abs/2407.21783
flash-attention— Tri Dao. github.com/Dao-AILab/flash-attention