Compression and efficiency optimization¶
Goal¶
Make the model cheaper and faster to serve without losing too much quality.
Inference cost dominates total LLM economics for any model that gets non-trivial usage. A 5% latency improvement at scale can dwarf the cost of post-training.
Methods¶
Quantization¶
Reduce precision.
Common formats:
- FP16 / BF16 — half-precision; standard.
- FP8 — newer (H100+); ~2× speedup over BF16 with minor quality loss.
- INT8 — integer 8-bit quantization; popular for serving.
- INT4 — 4-bit; aggressive quality/cost tradeoff. Used heavily for local inference.
- Weight-only quantization — quantize weights, keep activations in higher precision (GPTQ, AWQ, GGUF).
- Activation-aware quantization — preserves outliers (AWQ, SmoothQuant).
| Format | Memory savings | Typical quality impact |
|---|---|---|
| BF16 → FP8 | 2× | Negligible |
| BF16 → INT8 | 2× | Small |
| BF16 → INT4 (AWQ/GPTQ) | 4× | Modest, task-dependent |
| BF16 → INT4 (GGUF Q4_K_M) | 4× | Modest, task-dependent |
| Below INT4 | 5×+ | Significant; only for narrow tasks |
Tradeoffs:
- Lower memory and faster inference
- Possible quality degradation (especially on math, code, low-resource languages)
- Harder for small models — quantization noise hurts more relatively
- Calibration sets and quantization-aware training help
AWQ vs GPTQ vs GGUF
- AWQ — best quality at INT4 for serving on GPU.
- GPTQ — well-supported across frameworks.
- GGUF —
llama.cppformat. Best for CPU / Apple Silicon / consumer GPUs.
Distillation¶
Train a smaller student model to imitate a larger teacher.
Data:
- Teacher responses
- Teacher rationales (chain of thought)
- Preference-selected teacher outputs
- Tool trajectories
- Domain examples
- Logits from the teacher (for token-level distillation)
Modern open recipes use distillation aggressively. DeepSeek-R1 distillation produced strong small reasoning models from the large RL-trained teacher; this pattern is now common.
Pruning¶
Remove weights, heads, neurons, or layers.
Less common for frontier LLMs than quantization/distillation. Works better with structured pruning (whole heads or layers) than unstructured.
Speculative decoding¶
Use a smaller draft model to propose tokens; the larger model verifies in a single forward pass.
When the draft and target agree, you get multiple tokens per forward pass — a 2–3× throughput improvement is typical without quality loss.
Variants:
- Vanilla speculative decoding — Leviathan et al., 2022. arxiv.org/abs/2211.17192
- Medusa — multi-head prediction; no separate draft model. arxiv.org/abs/2401.10774
- EAGLE — feature-level speculation. arxiv.org/abs/2401.15077
- Lookahead decoding — n-gram-based.
KV-cache optimization¶
Important for long-context serving. Often the dominant memory consumer at long contexts.
Methods:
- Paged attention — vLLM. Treat KV cache like virtual memory; eliminate fragmentation. arxiv.org/abs/2309.06180
- Quantized KV cache — INT8 / INT4 KV cache; halves or quarters KV memory.
- Sliding window cache — discard older tokens past a window.
- Prefix caching — reuse KV cache across requests sharing a prefix.
- Cache eviction — H2O, StreamingLLM-style attention-score-based eviction.
- Attention sinks — keep the first few tokens always in cache to maintain stability.
Parameter-efficient fine-tuning¶
For task-specific adaptation without retraining the whole model:
- LoRA — low-rank adapters injected into linear layers. arxiv.org/abs/2106.09685
- QLoRA — LoRA on top of a 4-bit quantized base; very memory-efficient. arxiv.org/abs/2305.14314
- Adapters — small residual modules.
- Prefix tuning — learn prefix tokens.
- Prompt tuning — learn soft prompt embeddings.
Useful for task-specific adaptation, less often for full frontier post-training. LoRA dominates this category in 2025.
Practical tips¶
- Measure before optimizing. Profile latency, KV-cache memory, and throughput on real production traces. The bottleneck is often not where you'd guess.
- Don't quantize and pray. Run a full eval suite (capability + safety + format) on the quantized model. Quality regressions are real and uneven.
- Keep a higher-precision fallback. For high-stakes queries, route to the unquantized model.
- Prefix-cache aggressively. A few percent of system-prompt overlap unlocks large savings.
- Speculative decoding pays off most for long outputs. Short outputs see less benefit relative to its overhead.
Further reading¶
- Frantar et al., "GPTQ", 2022. arxiv.org/abs/2210.17323
- Lin et al., "AWQ", 2023. arxiv.org/abs/2306.00978
- Hu et al., "LoRA", 2021. arxiv.org/abs/2106.09685
- Dettmers et al., "QLoRA", 2023. arxiv.org/abs/2305.14314
- Leviathan et al., "Speculative Decoding", 2022. arxiv.org/abs/2211.17192
- Kwon et al., "PagedAttention / vLLM", 2023. arxiv.org/abs/2309.06180
- vLLM — github.com/vllm-project/vllm
- SGLang — github.com/sgl-project/sglang
- TensorRT-LLM — NVIDIA. github.com/NVIDIA/TensorRT-LLM
llama.cpp— github.com/ggerganov/llama.cppbitsandbytes— quantization library. github.com/bitsandbytes-foundation/bitsandbytes