Data collection and preparation¶
Goal¶
Create a massive, diverse, legally usable, high-quality training corpus that teaches language, reasoning, code, facts, style, multilingual ability, and domain knowledge.
Pretraining data is usually the most important asset. The LIMA paper is a useful reminder: much of an LLM's knowledge and capability comes from pretraining, while alignment data often teaches behavior and format rather than adding broad factual knowledge.
Common data sources¶
Web text¶
This is usually the largest source.
Examples:
- Common Crawl-style web snapshots
- RefinedWeb-like filtered web corpora
- News pages
- Blogs
- Documentation
- Forums
- Educational sites
- Public-domain or licensed text
Web data gives breadth, but it is noisy. Raw web text contains spam, boilerplate, duplicate pages, SEO sludge, adult content, hate, scams, malware snippets, and benchmark contamination.
Reference web corpora
- C4 — cleaned Common Crawl from the T5 paper. huggingface.co/datasets/allenai/c4
- RefinedWeb — Falcon's filtered web corpus. arxiv.org/abs/2306.01116
- FineWeb / FineWeb-Edu — HuggingFace's heavily-filtered web corpus, with an "educational quality" subset. huggingface.co/datasets/HuggingFaceFW/fineweb
- Dolma — AI2's open pretraining corpus used for OLMo. huggingface.co/datasets/allenai/dolma
- The Pile — EleutherAI's classic 800 GiB diverse corpus. arxiv.org/abs/2101.00027
Books and long-form text¶
Books help with:
- Coherence over long spans
- Narrative structure
- Style
- Domain depth
- Reasoning over paragraphs and chapters
Sources may include public-domain books, licensed book corpora, textbooks, manuals, and long-form essays.
Licensing
Books are the highest-leverage and highest-legal-risk data source. The state of the art in 2025 still relies heavily on disputed corpora; do not assume because a corpus is on HuggingFace it is safe to train on commercially. Consult counsel.
Code¶
Code data is crucial for modern LLMs, even if the target product is not a coding model.
Useful sources:
- Public repositories
- Documentation
- Code comments
- Issue discussions
- Pull request discussions
- Stack Overflow-like Q&A
- API examples
- Unit tests
- Notebooks
Code improves:
- Symbolic reasoning
- Tool use
- Structured output
- Exact formatting
- Multi-step planning
- Debugging
- Following formal constraints
Reference code corpora
- The Stack v2 — 67 TB of permissively-licensed code. huggingface.co/datasets/bigcode/the-stack-v2
- CodeParrot, CodeContests, APPS — smaller curated code datasets.
- StarCoder training data — see arxiv.org/abs/2305.06161.
Academic and technical text¶
Examples:
- arXiv papers
- PubMed abstracts or licensed biomedical text
- Legal documents
- Patents
- Standards
- Scientific documentation
- Mathematical problem datasets
This data helps with technical vocabulary, long-range reasoning, and expert-domain competence.
Reference academic corpora
- arXiv bulk — full-text via Kaggle or arXiv's S3 export.
- PubMed Central Open Access Subset — publicly licensed biomedical full text.
- OpenWebMath / proof-pile — math-heavy open corpora.
Q&A and conversation data¶
Examples:
- Public Q&A forums
- Help forums
- Dialogue datasets
- Customer support-style interactions
- Human-written instruction datasets
- Synthetic instruction data
Raw Q&A data is useful, but needs heavy filtering because online answers vary wildly in correctness.
Multilingual data¶
A state-of-the-art model needs intentional multilingual coverage, not just incidental non-English web text.
Important considerations:
- Language balance
- Script coverage
- Low-resource language quality
- Translationese detection
- Toxicity and spam filtering per language
- Tokenizer efficiency across scripts
- Cultural and dialect coverage
Reference multilingual corpora
- mC4 — multilingual C4.
- CulturaX — cleaned multilingual web corpus, 167 languages. huggingface.co/datasets/uonlp/CulturaX
- HPLT — High Performance Language Technologies multilingual data.
- OSCAR — large multilingual web corpus with per-language filters.
Synthetic data¶
Synthetic data is now central to high-performing post-training and often used in pretraining or continued pretraining.
Examples:
- Model-generated instructions
- Model-generated answers
- Reasoning traces
- Code problems
- Unit tests
- Tool-use trajectories
- Function-calling examples
- Long-context question-answer pairs
- Safety refusals and safe completions
- Data rewritten for style, clarity, or difficulty
Synthetic data can be extremely effective, but it risks model collapse, overfitting to the teacher model's style, hidden hallucinations, and loss of diversity.
Synthetic data hygiene
- Mix synthetic with human-written to anchor diversity.
- Verify synthetic answers when possible (unit tests, math checkers, schema validators).
- Watch for teacher-style imprinting — verbose disclaimers, signature openings, etc.
- Track which model produced which synthetic data; you may need to regenerate when the teacher updates.
Practical tips¶
- Build a data lake first, then a corpus. Keep raw, cleaned, deduplicated, filtered, and final mixture as separate stages. Versioning each stage saves enormous time during ablation.
- Score, don't just filter. Quality filters are easier to tune as continuous scores you threshold later than as hard binary filters baked into the pipeline.
- Sample and read. No tooling replaces reading 500 random documents from your final corpus. You will find issues.
- Document the mixture. A pretraining run is reproducible only if the mixture (and its proportions, and the seeds for sampling) is documented.
Further reading¶
- The Pile — Gao et al., 2020. arxiv.org/abs/2101.00027
- RefinedWeb — Penedo et al., 2023. The case for "the web is enough, if you filter it." arxiv.org/abs/2306.01116
- FineWeb / FineWeb-Edu — HuggingFace, 2024. The current open-corpus state of the art. huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1
- Dolma — Soldaini et al., 2024. AI2's open pretraining corpus. arxiv.org/abs/2402.00159
- OLMo 2 — AI2's open recipe with full data documentation. arxiv.org/abs/2501.00656
- Llama 3 paper, §3 (data) — describes a modern frontier-grade data pipeline. arxiv.org/abs/2407.21783
- DataComp-LM — large-scale data-centric benchmark. arxiv.org/abs/2406.11794
- CulturaX — Nguyen et al., 2023. arxiv.org/abs/2309.09400