Deduplication stops a model from memorizing the same example twice and then looking smart on a leaked validation row. I compare records by stable identifiers, normalized text, and near-duplicate hashes, not only by file names. Deduplication sits after cleaning and before the train / validation / test split.
I do not delete the only copy of a rare class because two rows share a prefix. The policy is written: exact match, normalized match, or a similarity threshold you can explain. For images, a perceptual hash catches a resized copy. For text, a normalized hash catches repeated boilerplate. Junk is a different pass: garbage is not the same as a duplicate of a good row. See junk removal.
Acceptance is a before-and-after count, a sample of removed pairs, and a check that no validation example has an identical twin in train. Removed keys stay in a side file. A single-source pass is often 1–2 weeks once the schema exists.
