AI Dataset Engineering

Train, validation, and test split

A train / validation / test split is honest only when the same user, document, or photo cannot sit on both sides.

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

A train / validation / test split is honest only when the same user, document, or photo cannot sit on both sides. I group by the unit that would leak: account, ticket thread, source file, near-duplicate cluster from deduplication. Train, validation, and test are frozen files, not a fresh random cut every run.

Class balance is checked, not wished for. A rare class is stratified when the counts allow it; when they do not, the report says the test slice is too small to trust. Time-based tasks split by time, not by a shuffle that lets next week train on yesterday. Synthetic twins stay with their real parent on one side.

Acceptance is the split manifest, the grouping key, counts per slice and class, and a leakage check that fails if a normalized twin crosses the line. Writing the split for an existing clean corpus is often 1–2 weeks. The numbers are then part of dataset quality control.

Acceptance criteria

Done when

  • Schema and field meanings are written down
  • Train / validation / test split is reproducible and checked for leakage
  • Quality report lists counts, removed duplicates, and known gaps

Deliverables

  • Dataset files in the agreed format
  • Reproducible preparation script
  • Quality report

Out of scope

  • Training the model and production deployment
  • Legal opinion on personal data and third-party licenses
  • Annotator volume beyond the agreed sample unless it is in the quote

The final acceptance checklist is confirmed in the brief or contract; the list above is a scope alignment guide.

Ballpark estimate

Scope size
Extras

FAQ

Tap a question to expand the answer.

What ratios do you use?

Often 80/10/10, but the grouping key wins over a pretty ratio. I will not break a document across slices to hit a percentage.

Can we reshuffle later?

Only with a new version of the dataset. The accepted split stays fixed so scores stay comparable.

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

View full service page