AI Dataset Engineering

Deduplication

Deduplication stops a model from memorizing the same example twice and then looking smart on a leaked validation row.

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

Deduplication stops a model from memorizing the same example twice and then looking smart on a leaked validation row. I compare records by stable identifiers, normalized text, and near-duplicate hashes, not only by file names. Deduplication sits after cleaning and before the train / validation / test split.

I do not delete the only copy of a rare class because two rows share a prefix. The policy is written: exact match, normalized match, or a similarity threshold you can explain. For images, a perceptual hash catches a resized copy. For text, a normalized hash catches repeated boilerplate. Junk is a different pass: garbage is not the same as a duplicate of a good row. See junk removal.

Acceptance is a before-and-after count, a sample of removed pairs, and a check that no validation example has an identical twin in train. Removed keys stay in a side file. A single-source pass is often 1–2 weeks once the schema exists.

Acceptance criteria

Done when

  • Schema and field meanings are written down
  • Train / validation / test split is reproducible and checked for leakage
  • Quality report lists counts, removed duplicates, and known gaps

Deliverables

  • Dataset files in the agreed format
  • Reproducible preparation script
  • Quality report

Out of scope

  • Training the model and production deployment
  • Legal opinion on personal data and third-party licenses
  • Annotator volume beyond the agreed sample unless it is in the quote

The final acceptance checklist is confirmed in the brief or contract; the list above is a scope alignment guide.

Ballpark estimate

Scope size
Extras

FAQ

Tap a question to expand the answer.

Exact match or near-duplicate?

Exact match is the default. Near-duplicates need a threshold you can defend, because aggressive hashing drops real variants.

Do you keep an audit of removed rows?

Yes. Removed keys stay beside the dataset so a later review can see what left and why.

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

View full service page