Data¶
Pretraining data is usually the most important asset in an LLM project. The LIMA paper is a useful reminder: much of an LLM's knowledge and capability comes from pretraining, while alignment data often teaches behavior and format rather than adding broad factual knowledge.
This section covers the full data lifecycle:
- Collection & sources — web, books, code, math, multilingual, synthetic. What each source contributes and what to watch for.
- Cleaning & filtering — deduplication, quality filtering, safety filtering, PII removal, benchmark decontamination.
For SFT and preference data specifically, see:
- Supervised fine-tuning — instruction-tuning data categories and mixes.
- Preference data — sources, formats, and verifiable rewards.
The cardinal rules¶
- Quality beats quantity after a point. Filter aggressively. A smaller, cleaner corpus often beats a larger, dirtier one.
- Diversity is a real signal. A mixture covering web, books, code, math, multilingual, and academic text outperforms any one source alone.
- Decontamination matters. Benchmark contamination silently inflates scores. Run n-gram and embedding-based decontamination before declaring a result.
- PII filtering is imperfect. Plan for memorization audits and red-team probes after training.
- Document provenance. For every shard, track source, license, date, and any transformations applied. You will need this for legal and reproducibility reasons.