Skip to content

Context length extension

Goal

Increase the model's usable context window.

Examples:

  • 4K → 8K
  • 8K → 32K
  • 32K → 128K
  • 128K → 1M in specialized systems

Llama 3's largest model is reported with context up to 128K tokens. Gemini 1.5 and several research models reach 1M+. The frontier keeps moving.

Why long context is hard

Attention is expensive. Vanilla self-attention is quadratic in sequence length. Doubling context can roughly quadruple attention cost.

FlashAttention (Dao et al., 2022) helps by computing exact attention with fewer high-bandwidth memory reads/writes, improving speed and memory efficiency. It does not change asymptotic complexity but slashes the constant factor and enables much longer sequences in practice.

KV-cache memory at inference is the second bottleneck. At 128K context, KV cache for a 70B-class model can be tens of gigabytes per request — see Compression for mitigations (paged attention, quantized KV, GQA).

Main approaches

1. Train from scratch at long context

Best quality, very expensive.

Pros:

  • Native long-context ability
  • Fewer extrapolation issues

Cons:

  • High compute cost
  • Harder batching
  • More memory pressure

Few labs do this end-to-end; most bootstrap from a shorter-context base model and extend.

2. RoPE scaling / interpolation

For RoPE-based models, modify positional frequency scaling so the model can operate beyond its original context.

Methods include:

  • Linear (Position) Interpolation — Chen et al., 2023. Compresses positions linearly into the original range. Simple, works decently. arxiv.org/abs/2306.15595
  • NTK-aware scaling — adjusts the RoPE base θ to interpolate high-frequency dimensions less. Works well with no fine-tuning for moderate extensions.
  • YaRN — Peng et al., 2023. Dimension-dependent scaling. State-of-the-art with substantially fewer training tokens. arxiv.org/abs/2309.00071
  • LongRoPE — Microsoft, 2024. Pushes to 2M+ context. arxiv.org/abs/2402.13753

YaRN specifically targets efficient context-window extension for RoPE models and reports requiring substantially fewer tokens and training steps than previous methods.

Default starting point

Pick YaRN unless you have a specific reason not to. It is the most-cited, best-supported method in 2025 open-weights work, and major inference frameworks (vLLM, SGLang, llama.cpp) implement it natively.

3. Continued pretraining on long sequences

After changing the positional scheme, train on long-context data.

Data examples:

  • Books
  • Long documents
  • Legal contracts
  • Code repositories
  • Multi-file code tasks
  • Long conversations
  • Retrieval-augmented synthetic tasks
  • Needle-in-a-haystack examples
  • Long-context QA

Token budget for the long-context phase is typically small — billions of tokens, not trillions — but with careful mixing.

4. Architectural changes

Examples:

  • Sliding-window attention — Mistral 7B v0.1 used 4K window with rolling buffer.
  • Global tokens — a few special tokens that attend everywhere.
  • Sparse attention — block-sparse, longformer-style.
  • Memory tokens — learned compressed representations.
  • Recurrence — Transformer-XL, RWKV-style.
  • State-space hybrids — Mamba, Jamba, RecurrentGemma.
  • Retrieval-augmented generation — see RAG.
  • External memory — explicit retrieval over past content.

Long-context training data

You need tasks that require using far-away information.

Bad long-context data:

  • Just concatenated random documents
  • Long text where the answer is local
  • Repetitive filler
  • Synthetic "needle" tasks only

Good long-context data:

  • Multi-section document QA
  • Cross-document synthesis
  • Long codebase navigation
  • Legal/clinical/scientific evidence aggregation
  • Conversation memory tasks
  • Temporal consistency tasks
  • Multi-hop retrieval within context

Needle-only training is a trap

A model trained only on synthetic needle-in-a-haystack examples can ace needle benchmarks while still failing real long-context tasks. The reasoning skills that matter (multi-hop synthesis, cross-section consistency, "the answer is implied across three places") need explicit training data.

Long-context evals

Evaluate:

  • Needle-in-a-haystack
  • Multi-needle retrieval
  • LongBench-style tasks
  • Book QA
  • Long summarization
  • Repository-level code understanding
  • Long conversation consistency
  • Lost-in-the-middle behavior

Important: passing needle retrieval does not mean the model can reason well over long context.

Eval suites worth running

Practical tips

  • Don't trust nominal context length. Many models claim 128K but degrade noticeably past 32K. Run RULER or LongBench, not just needle.
  • Test "lost in the middle". Models often answer well from the start and end of the context but miss information from the middle. See Liu et al., 2023.
  • Tune the inference stack. Long context exposes serving bugs (KV cache eviction, attention numerical issues, prefix caching). Test the served model, not just training-time behavior.
  • Mix short and long during training. Pure long-context fine-tuning regresses short-context behavior; a mix of sequence lengths is essential.

Further reading