Junk removal takes out rows that should never teach a model: empty files, repeated headers, OCR garbage, spam, placeholder “test” records, and boilerplate that appears on every page. Junk removal is not deduplication. A duplicate of a good example is still a good example; junk has no label worth learning.
Rules are named before they run. “Too short”, “mostly symbols”, “known test account”, “template footer”. I review a sample of what the rule would drop, including a few rows it should keep, so a rare class is not swept out with the trash. Personal data that must leave the corpus is masked or dropped under the brief, not mixed into a vague “clean it up”.
Acceptance is the rule list, the drop counts, and a signed sample of removed and kept rows. The side file of removed keys stays with the dataset. A rules pass on one text source is often 1–2 weeks. Quality numbers come later in dataset quality control.
