Skip to content

Multimodal extension

Not every LLM is multimodal, but many state-of-the-art assistants are. Vision is the most common addition; audio and video are growing.

Common stages

Vision encoder pretraining

Use a vision encoder such as a ViT-like model. Most modern open multimodal LLMs use a frozen pretrained vision encoder (CLIP, SigLIP, OpenCLIP) rather than training one from scratch.

Projector training

Train a small adapter/projector from image embeddings to the LLM embedding space. This is typically a 2-layer MLP or a Q-Former.

This is the cheapest stage — only the projector trains. The vision encoder and the LLM backbone are frozen. Done first because it's fast and sets up the next stage.

Multimodal instruction tuning

Now unfreeze (some of) the LLM and train end-to-end on multimodal instruction data:

  • Image captioning
  • Visual QA
  • OCR
  • Chart QA
  • Document QA
  • UI screenshots
  • Grounding (point to or box specific objects)
  • Spatial reasoning
  • Medical images, if domain-specific and licensed

Preference optimization

Compare multimodal answers — same DPO/PPO machinery as text, with the prompt now including images.

Safety

Train for:

  • Face recognition limits
  • Medical image caution (don't replace a doctor)
  • Sensitive attribute inference limits (don't infer race / gender / health from a face)
  • OCR privacy
  • Sexual / violent image handling

Failure modes

  • Hallucinated visual details — describing things not in the image.
  • Poor counting — "how many people are in this photo" is famously hard.
  • Weak spatial reasoning — left/right/above/below errors.
  • OCR errors — especially handwriting, low-resolution, or unusual fonts.
  • Chart misreading — reading wrong axes, missing labels.
  • Overconfidence — declaring "this X-ray shows pneumonia" when uncertain.
  • Privacy issues — identifying people, inferring sensitive attributes.
  • Image-prompt-injection — text inside an image overriding instructions.

Reference open multimodal recipes

  • LLaVA — Liu et al. The most-cited open recipe (CLIP + projector + LLM). arxiv.org/abs/2304.08485 and arxiv.org/abs/2310.03744
  • MiniGPT-4 / MiniGPT-v2 — early open multimodal recipes.
  • Qwen-VL / Qwen2-VL — Alibaba. Strong open multimodal models with detailed reports. arxiv.org/abs/2308.12966
  • Idefics2 — HuggingFace. Open multimodal model and recipe. arxiv.org/abs/2405.02246
  • Llama 3.2 Vision / Llama 4 — Meta. Closed-recipe but open-weights.
  • InternVL — multimodal model series with strong document understanding.

Practical tips

  • Start with a frozen vision encoder. Training your own ViT from scratch is rarely worth it. CLIP or SigLIP works for most use cases.
  • Match resolution to the task. Document understanding needs higher resolution (1024px+) than general scene description.
  • Train on diverse aspect ratios. Square-cropped training hurts performance on portrait/landscape inputs.
  • Add OCR data explicitly. Implicit OCR via image-caption pairs is much weaker than explicit OCR training.
  • Don't forget text-only data. Pure multimodal training regresses pure-text performance. Mix text-only chat data through multimodal SFT.
  • Test on red-team images. Adversarial images with embedded text instructions are the multimodal version of prompt injection.

Further reading