A train / validation / test split is honest only when the same user, document, or photo cannot sit on both sides. I group by the unit that would leak: account, ticket thread, source file, near-duplicate cluster from deduplication. Train, validation, and test are frozen files, not a fresh random cut every run.
Class balance is checked, not wished for. A rare class is stratified when the counts allow it; when they do not, the report says the test slice is too small to trust. Time-based tasks split by time, not by a shuffle that lets next week train on yesterday. Synthetic twins stay with their real parent on one side.
Acceptance is the split manifest, the grouping key, counts per slice and class, and a leakage check that fails if a normalized twin crosses the line. Writing the split for an existing clean corpus is often 1–2 weeks. The numbers are then part of dataset quality control.
