Data cleaning and filtering¶
Deduplication¶
Deduplication removes repeated documents, paragraphs, or spans.
Why it matters:
- Prevents memorization
- Improves effective data diversity
- Reduces train/test contamination
- Prevents common templates from dominating
- Reduces wasted compute
Typical levels:
- Exact URL/document deduplication
- Near-duplicate document deduplication
- Paragraph-level deduplication
- MinHash / SimHash clustering
- Code clone detection
- Benchmark decontamination
Deduplication tooling
datasketch— Python MinHash, MinHashLSH. github.com/ekzhu/datasketchtext-dedup— fast MinHash+LSH for text corpora. github.com/ChenghaoMou/text-dedupdatatrove— HuggingFace's distributed data-processing pipeline used to build FineWeb. github.com/huggingface/datatrove
The classic reference for why deduplication matters: Lee et al., "Deduplicating Training Data Makes Language Models Better", 2021. arxiv.org/abs/2107.06499
Quality filtering¶
Quality filters attempt to keep high-signal text.
Signals include:
- Language identification confidence
- Perplexity under a reference model
- Length thresholds
- Character entropy
- Ratio of alphabetic to non-alphabetic characters
- Boilerplate density
- Repetition
- Link density
- Markdown or HTML artifacts
- Toxicity scores
- Spam classifiers
- Educational-value classifiers
- Domain-specific quality classifiers
A common pattern is to train small classifiers that distinguish "high-quality reference text" from random web text, then score the full corpus.
FineWeb-Edu's quality classifier
HuggingFace trained a small classifier to predict whether a piece of web text was educationally valuable, using ~500K Llama-3-rated examples. They then re-scored 15T+ tokens of FineWeb and kept the top fraction. Result: better downstream evals at a fraction of the data. See the FineWeb blog post.
Safety filtering¶
Pretraining does not usually remove every unsafe concept, because the model needs to understand the world. But most pipelines filter or downweight:
- Malware instructions
- Phishing kits
- Personal data
- Explicit personal records
- Extreme violence
- Sexual abuse material
- Spam
- Weapon construction instructions
- Hate and harassment
- Fraud instructions
- Low-quality medical or legal advice
The tradeoff is subtle: over-filtering can make the model ignorant or brittle; under-filtering increases misuse and toxic behavior.
PII removal¶
Common steps:
- Detect emails, phone numbers, addresses, social security numbers, access tokens, API keys
- Remove or mask personal records
- Use regexes plus learned detectors
- Run secret scanners on code
- Remove private keys and credentials
PII filtering is imperfect. This is why later safety evaluation and memorization testing are important.
Memorization probes
After training, sample random spans from the training corpus and prompt the model with the prefix. Measure how often the model reproduces the suffix verbatim. High reproduction rates on PII-bearing spans are a deployment blocker. See Carlini et al., "Quantifying Memorization Across Neural Language Models", arxiv.org/abs/2202.07646.
Benchmark decontamination¶
You remove or mark examples overlapping with evaluation benchmarks.
Methods:
- Exact match
- n-gram overlap (commonly 13-gram)
- MinHash similarity
- Embedding similarity
- Problem-statement matching
- Code benchmark duplicate detection
- Manual review for high-value evals
Without decontamination, benchmark results can be inflated.
Common decontamination mistakes
- Only checking the test set, not the training prompts of leaked instruction datasets.
- Using a single n-gram length (13-gram is common but easily evaded by paraphrase).
- Not decontaminating post-training data — SFT and preference data also leak.
- Treating decontamination as a one-time pass rather than a continuous practice.
Practical tips¶
- Run filters in cheapest-to-most-expensive order. Length thresholds and language ID first, perplexity and classifier-based scoring last. You will discard 60–90% of raw web text before the expensive scorers see it.
- Keep the rejected slices. They are useful for ablations, for training spam classifiers, and for safety probes.
- Score per-shard. Random sampling across shards uncovers shard-specific bugs (encoding issues, wrong language ID, missing fields) faster than aggregate metrics.
- Sanity-check final mixture proportions. If your "code" slice is suddenly 70% of the corpus, something upstream changed.
Further reading¶
- Lee et al., "Deduplicating Training Data Makes Language Models Better", 2021. arxiv.org/abs/2107.06499
- Carlini et al., "Quantifying Memorization Across Neural Language Models", 2022. arxiv.org/abs/2202.07646
- Penedo et al., "The RefinedWeb Dataset for Falcon LLM", 2023. arxiv.org/abs/2306.01116
- Soldaini et al., "Dolma", 2024. arxiv.org/abs/2402.00159
- Albalak et al., "A Survey on Data Selection for Language Models", 2024. arxiv.org/abs/2402.16827
- HuggingFace
datatrove— production-grade data-processing pipeline. github.com/huggingface/datatrove presidio— Microsoft's open-source PII detection toolkit. github.com/microsoft/presidio