Base model architecture¶
Most state-of-the-art LLMs are decoder-only Transformers trained with next-token prediction. The architecture has converged remarkably across labs since 2022 — the open-weights Llama, Mistral, Qwen, Gemma, and DeepSeek families share most of their architectural choices.
Core components¶
Token embeddings¶
Map token IDs to vectors of size d_model. The output projection (the "unembedding") is sometimes tied to the input embedding, sometimes separate. Most modern open models do not tie weights, since untied weights consistently win at scale.
Transformer blocks¶
Each block usually contains:
- Self-attention
- MLP / feed-forward network
- Residual connections
- Layer normalization or RMSNorm
- Positional encoding
- Sometimes attention or MLP variants
The residual stream stays in d_model dimensions throughout the network. Attention and MLP both read from and write to the residual stream.
Attention¶
The model attends to previous tokens. In a causal decoder, token t can see tokens ≤ t, but not future tokens.
Common attention variants:
- Multi-head attention (MHA) — original Transformer design. Each head has its own Q, K, V.
- Multi-query attention (MQA) — all heads share K and V projections. Cheaper inference, slight quality loss.
- Grouped-query attention (GQA) — middle ground; groups of heads share K/V. The current default for most open models.
- Sliding-window attention — each token only attends to a fixed-size local window.
- Local/global attention hybrids — most layers local, a few global.
- Sparse attention variants — block-sparse, longformer-style, etc.
GQA is popular because it reduces KV-cache memory during inference, which dominates serving cost at long contexts.
Positional embeddings¶
Common choices:
- RoPE / rotary embeddings — rotates Q and K vectors by position-dependent angles. Dominant in modern open LLMs because it generalizes to longer contexts (with scaling).
- ALiBi — adds a position-dependent bias to attention scores. Simple, but less common in 2025.
- Learned positional embeddings — fixed-length, doesn't extrapolate; older models.
- Relative position schemes — T5-style biases, mostly historical now.
RoPE is widely used in open LLMs and is central to many context-extension methods.
MLP block¶
Often the largest compute component. Usually up_proj, gate_proj, down_proj with a hidden dimension ~4× d_model (for SwiGLU, typically ~2.7× to keep parameter count similar).
Common variants:
- SwiGLU — gated linear unit with Swish activation. Current default.
- GeGLU — gated linear unit with GELU activation.
- ReLU/GELU older variants — standard 2-matrix MLPs without gating; mostly historical.
Normalization¶
RMSNorm is common in modern open models. It is simpler and slightly faster than LayerNorm, with no measurable quality loss.
Placement:
- Pre-norm (norm before attention/MLP) — standard. Stabilizes training at depth.
- Post-norm — original Transformer design; harder to train at scale.
Dense vs Mixture-of-Experts¶
Dense models activate all parameters per token.
Mixture-of-Experts (MoE) models route each token to a subset of expert networks (e.g., 2 of 8 experts active per token).
MoE advantages:
- More total parameters for similar inference FLOPs
- Better specialization
- Strong scaling
MoE difficulties:
- Routing instability
- Load balancing
- Serving complexity
- Expert parallelism
- More complicated fine-tuning
When MoE makes sense
MoE pays off when you have the serving infrastructure (expert parallelism) and your bottleneck is quality at fixed FLOPs. If your bottleneck is memory at fixed cost, dense is often easier. Mixtral, DBRX, DeepSeek-V3, and Llama 4 all use MoE.
A representative modern recipe¶
A typical 7B-class open model in 2025 looks something like:
| Choice | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Layers | 32 |
d_model |
4096 |
| Attention heads | 32 |
| GQA groups | 8 (KV heads) |
| MLP hidden | 14336 (SwiGLU) |
| Normalization | RMSNorm, pre-norm |
| Position | RoPE, base θ = 500,000 |
| Activation | SwiGLU |
| Vocab | 128K (byte-level BPE) |
| Context (pretraining) | 8K → extended |
| Precision | BF16 weights, FP32 accumulation |
This is approximately the Llama 3 8B / Qwen 2.5 7B shape. Larger models scale layers, d_model, and heads roughly together.
Practical tips¶
- Don't reinvent. Start from a known-good recipe (Llama 3, Qwen 2.5, OLMo 2). Architectural novelty is rarely the limiting factor at small scale and often regresses at large scale.
- Match width and depth to memory. Tensor parallelism efficiency depends on
d_modelbeing divisible by the TP degree; pickd_modelaccordingly. - Reserve KV-cache budget. For long-context serving, KV-cache memory often dominates. GQA reduces this; MQA reduces it further but can hurt quality.
- Test with a small model first. Most architecture bugs show up at 100M–1B scale and are cheap to find there.
Further reading¶
- Vaswani et al., "Attention Is All You Need", 2017. The original Transformer. arxiv.org/abs/1706.03762
- Su et al., "RoFormer: Enhanced Transformer with Rotary Position Embedding", 2021. RoPE. arxiv.org/abs/2104.09864
- Shazeer, "GLU Variants Improve Transformer", 2020. SwiGLU/GeGLU. arxiv.org/abs/2002.05202
- Zhang & Sennrich, "Root Mean Square Layer Normalization", 2019. RMSNorm. arxiv.org/abs/1910.07467
- Ainslie et al., "GQA: Training Generalized Multi-Query Transformer Models", 2023. arxiv.org/abs/2305.13245
- Llama 3 paper — Meta, 2024. arxiv.org/abs/2407.21783
- OLMo 2 paper — AI2, 2024. arxiv.org/abs/2501.00656
- Mixtral 8×7B paper — Mistral AI, 2024. The reference open-weights MoE. arxiv.org/abs/2401.04088
- DeepSeek-V3 technical report, 2024. State-of-the-art open MoE. arxiv.org/abs/2412.19437
nanoGPT— Karpathy's minimal Transformer training code. github.com/karpathy/nanoGPTllm.c— Karpathy's pure-CUDA Transformer training. github.com/karpathy/llm.c- The Annotated Transformer — Harvard NLP. Code-by-code walkthrough. nlp.seas.harvard.edu/annotated-transformer