Skip to content

Base model architecture

Most state-of-the-art LLMs are decoder-only Transformers trained with next-token prediction. The architecture has converged remarkably across labs since 2022 — the open-weights Llama, Mistral, Qwen, Gemma, and DeepSeek families share most of their architectural choices.

Core components

Token embeddings

Map token IDs to vectors of size d_model. The output projection (the "unembedding") is sometimes tied to the input embedding, sometimes separate. Most modern open models do not tie weights, since untied weights consistently win at scale.

Transformer blocks

Each block usually contains:

  • Self-attention
  • MLP / feed-forward network
  • Residual connections
  • Layer normalization or RMSNorm
  • Positional encoding
  • Sometimes attention or MLP variants

The residual stream stays in d_model dimensions throughout the network. Attention and MLP both read from and write to the residual stream.

Attention

The model attends to previous tokens. In a causal decoder, token t can see tokens ≤ t, but not future tokens.

Common attention variants:

  • Multi-head attention (MHA) — original Transformer design. Each head has its own Q, K, V.
  • Multi-query attention (MQA) — all heads share K and V projections. Cheaper inference, slight quality loss.
  • Grouped-query attention (GQA) — middle ground; groups of heads share K/V. The current default for most open models.
  • Sliding-window attention — each token only attends to a fixed-size local window.
  • Local/global attention hybrids — most layers local, a few global.
  • Sparse attention variants — block-sparse, longformer-style, etc.

GQA is popular because it reduces KV-cache memory during inference, which dominates serving cost at long contexts.

Positional embeddings

Common choices:

  • RoPE / rotary embeddings — rotates Q and K vectors by position-dependent angles. Dominant in modern open LLMs because it generalizes to longer contexts (with scaling).
  • ALiBi — adds a position-dependent bias to attention scores. Simple, but less common in 2025.
  • Learned positional embeddings — fixed-length, doesn't extrapolate; older models.
  • Relative position schemes — T5-style biases, mostly historical now.

RoPE is widely used in open LLMs and is central to many context-extension methods.

MLP block

Often the largest compute component. Usually up_proj, gate_proj, down_proj with a hidden dimension ~4× d_model (for SwiGLU, typically ~2.7× to keep parameter count similar).

Common variants:

  • SwiGLU — gated linear unit with Swish activation. Current default.
  • GeGLU — gated linear unit with GELU activation.
  • ReLU/GELU older variants — standard 2-matrix MLPs without gating; mostly historical.

Normalization

RMSNorm is common in modern open models. It is simpler and slightly faster than LayerNorm, with no measurable quality loss.

Placement:

  • Pre-norm (norm before attention/MLP) — standard. Stabilizes training at depth.
  • Post-norm — original Transformer design; harder to train at scale.

Dense vs Mixture-of-Experts

Dense models activate all parameters per token.

Mixture-of-Experts (MoE) models route each token to a subset of expert networks (e.g., 2 of 8 experts active per token).

MoE advantages:

  • More total parameters for similar inference FLOPs
  • Better specialization
  • Strong scaling

MoE difficulties:

  • Routing instability
  • Load balancing
  • Serving complexity
  • Expert parallelism
  • More complicated fine-tuning

When MoE makes sense

MoE pays off when you have the serving infrastructure (expert parallelism) and your bottleneck is quality at fixed FLOPs. If your bottleneck is memory at fixed cost, dense is often easier. Mixtral, DBRX, DeepSeek-V3, and Llama 4 all use MoE.

A representative modern recipe

A typical 7B-class open model in 2025 looks something like:

Choice Value
Architecture Decoder-only Transformer
Layers 32
d_model 4096
Attention heads 32
GQA groups 8 (KV heads)
MLP hidden 14336 (SwiGLU)
Normalization RMSNorm, pre-norm
Position RoPE, base θ = 500,000
Activation SwiGLU
Vocab 128K (byte-level BPE)
Context (pretraining) 8K → extended
Precision BF16 weights, FP32 accumulation

This is approximately the Llama 3 8B / Qwen 2.5 7B shape. Larger models scale layers, d_model, and heads roughly together.

Practical tips

  • Don't reinvent. Start from a known-good recipe (Llama 3, Qwen 2.5, OLMo 2). Architectural novelty is rarely the limiting factor at small scale and often regresses at large scale.
  • Match width and depth to memory. Tensor parallelism efficiency depends on d_model being divisible by the TP degree; pick d_model accordingly.
  • Reserve KV-cache budget. For long-context serving, KV-cache memory often dominates. GQA reduces this; MQA reduces it further but can hurt quality.
  • Test with a small model first. Most architecture bugs show up at 100M–1B scale and are cheap to find there.

Further reading