Skip to content

Compression and efficiency optimization

Goal

Make the model cheaper and faster to serve without losing too much quality.

Inference cost dominates total LLM economics for any model that gets non-trivial usage. A 5% latency improvement at scale can dwarf the cost of post-training.

Methods

Quantization

Reduce precision.

Common formats:

  • FP16 / BF16 — half-precision; standard.
  • FP8 — newer (H100+); ~2× speedup over BF16 with minor quality loss.
  • INT8 — integer 8-bit quantization; popular for serving.
  • INT4 — 4-bit; aggressive quality/cost tradeoff. Used heavily for local inference.
  • Weight-only quantization — quantize weights, keep activations in higher precision (GPTQ, AWQ, GGUF).
  • Activation-aware quantization — preserves outliers (AWQ, SmoothQuant).
Format Memory savings Typical quality impact
BF16 → FP8 2× Negligible
BF16 → INT8 2× Small
BF16 → INT4 (AWQ/GPTQ) 4× Modest, task-dependent
BF16 → INT4 (GGUF Q4_K_M) 4× Modest, task-dependent
Below INT4 5×+ Significant; only for narrow tasks

Tradeoffs:

  • Lower memory and faster inference
  • Possible quality degradation (especially on math, code, low-resource languages)
  • Harder for small models — quantization noise hurts more relatively
  • Calibration sets and quantization-aware training help

AWQ vs GPTQ vs GGUF

  • AWQ — best quality at INT4 for serving on GPU.
  • GPTQ — well-supported across frameworks.
  • GGUF — llama.cpp format. Best for CPU / Apple Silicon / consumer GPUs.

Distillation

Train a smaller student model to imitate a larger teacher.

Data:

  • Teacher responses
  • Teacher rationales (chain of thought)
  • Preference-selected teacher outputs
  • Tool trajectories
  • Domain examples
  • Logits from the teacher (for token-level distillation)

Modern open recipes use distillation aggressively. DeepSeek-R1 distillation produced strong small reasoning models from the large RL-trained teacher; this pattern is now common.

Pruning

Remove weights, heads, neurons, or layers.

Less common for frontier LLMs than quantization/distillation. Works better with structured pruning (whole heads or layers) than unstructured.

Speculative decoding

Use a smaller draft model to propose tokens; the larger model verifies in a single forward pass.

When the draft and target agree, you get multiple tokens per forward pass — a 2–3× throughput improvement is typical without quality loss.

Variants:

KV-cache optimization

Important for long-context serving. Often the dominant memory consumer at long contexts.

Methods:

  • Paged attention — vLLM. Treat KV cache like virtual memory; eliminate fragmentation. arxiv.org/abs/2309.06180
  • Quantized KV cache — INT8 / INT4 KV cache; halves or quarters KV memory.
  • Sliding window cache — discard older tokens past a window.
  • Prefix caching — reuse KV cache across requests sharing a prefix.
  • Cache eviction — H2O, StreamingLLM-style attention-score-based eviction.
  • Attention sinks — keep the first few tokens always in cache to maintain stability.

Parameter-efficient fine-tuning

For task-specific adaptation without retraining the whole model:

  • LoRA — low-rank adapters injected into linear layers. arxiv.org/abs/2106.09685
  • QLoRA — LoRA on top of a 4-bit quantized base; very memory-efficient. arxiv.org/abs/2305.14314
  • Adapters — small residual modules.
  • Prefix tuning — learn prefix tokens.
  • Prompt tuning — learn soft prompt embeddings.

Useful for task-specific adaptation, less often for full frontier post-training. LoRA dominates this category in 2025.

Practical tips

  • Measure before optimizing. Profile latency, KV-cache memory, and throughput on real production traces. The bottleneck is often not where you'd guess.
  • Don't quantize and pray. Run a full eval suite (capability + safety + format) on the quantized model. Quality regressions are real and uneven.
  • Keep a higher-precision fallback. For high-stakes queries, route to the unquantized model.
  • Prefix-cache aggressively. A few percent of system-prompt overlap unlocks large savings.
  • Speculative decoding pays off most for long outputs. Short outputs see less benefit relative to its overhead.

Further reading