Multimodal extension¶
Not every LLM is multimodal, but many state-of-the-art assistants are. Vision is the most common addition; audio and video are growing.
Common stages¶
Vision encoder pretraining¶
Use a vision encoder such as a ViT-like model. Most modern open multimodal LLMs use a frozen pretrained vision encoder (CLIP, SigLIP, OpenCLIP) rather than training one from scratch.
- CLIP — OpenAI. Original contrastively-trained vision-language encoder. arxiv.org/abs/2103.00020
- SigLIP — Google. Sigmoid loss; better than CLIP at small scale. arxiv.org/abs/2303.15343
- DINOv2 — Meta. Self-supervised, useful for dense visual understanding.
- OpenCLIP — open reproductions and training. github.com/mlfoundations/open_clip
Projector training¶
Train a small adapter/projector from image embeddings to the LLM embedding space. This is typically a 2-layer MLP or a Q-Former.
This is the cheapest stage — only the projector trains. The vision encoder and the LLM backbone are frozen. Done first because it's fast and sets up the next stage.
Multimodal instruction tuning¶
Now unfreeze (some of) the LLM and train end-to-end on multimodal instruction data:
- Image captioning
- Visual QA
- OCR
- Chart QA
- Document QA
- UI screenshots
- Grounding (point to or box specific objects)
- Spatial reasoning
- Medical images, if domain-specific and licensed
Preference optimization¶
Compare multimodal answers — same DPO/PPO machinery as text, with the prompt now including images.
Safety¶
Train for:
- Face recognition limits
- Medical image caution (don't replace a doctor)
- Sensitive attribute inference limits (don't infer race / gender / health from a face)
- OCR privacy
- Sexual / violent image handling
Failure modes¶
- Hallucinated visual details — describing things not in the image.
- Poor counting — "how many people are in this photo" is famously hard.
- Weak spatial reasoning — left/right/above/below errors.
- OCR errors — especially handwriting, low-resolution, or unusual fonts.
- Chart misreading — reading wrong axes, missing labels.
- Overconfidence — declaring "this X-ray shows pneumonia" when uncertain.
- Privacy issues — identifying people, inferring sensitive attributes.
- Image-prompt-injection — text inside an image overriding instructions.
Reference open multimodal recipes¶
- LLaVA — Liu et al. The most-cited open recipe (CLIP + projector + LLM). arxiv.org/abs/2304.08485 and arxiv.org/abs/2310.03744
- MiniGPT-4 / MiniGPT-v2 — early open multimodal recipes.
- Qwen-VL / Qwen2-VL — Alibaba. Strong open multimodal models with detailed reports. arxiv.org/abs/2308.12966
- Idefics2 — HuggingFace. Open multimodal model and recipe. arxiv.org/abs/2405.02246
- Llama 3.2 Vision / Llama 4 — Meta. Closed-recipe but open-weights.
- InternVL — multimodal model series with strong document understanding.
Practical tips¶
- Start with a frozen vision encoder. Training your own ViT from scratch is rarely worth it. CLIP or SigLIP works for most use cases.
- Match resolution to the task. Document understanding needs higher resolution (1024px+) than general scene description.
- Train on diverse aspect ratios. Square-cropped training hurts performance on portrait/landscape inputs.
- Add OCR data explicitly. Implicit OCR via image-caption pairs is much weaker than explicit OCR training.
- Don't forget text-only data. Pure multimodal training regresses pure-text performance. Mix text-only chat data through multimodal SFT.
- Test on red-team images. Adversarial images with embedded text instructions are the multimodal version of prompt injection.
Further reading¶
- LLaVA — Liu et al., 2023. arxiv.org/abs/2304.08485
- LLaVA-1.5 — Liu et al., 2023. arxiv.org/abs/2310.03744
- Qwen-VL — Bai et al., 2023. arxiv.org/abs/2308.12966
- Idefics2 — Laurençon et al., 2024. arxiv.org/abs/2405.02246
- Cambrian-1 — NYU. Vision-centric multimodal evaluation. arxiv.org/abs/2406.16860
- MMMU — multimodal benchmark suite. arxiv.org/abs/2311.16502
- MathVista — math+vision benchmark. arxiv.org/abs/2310.02255
- LMMs-Eval — open multimodal eval framework. github.com/EvolvingLMMs-Lab/lmms-eval