Dataset quality control is the report that says whether the corpus is allowed to train a model. I count rows, classes, empty fields, duplicates removed, junk removed, and synthetic share. Quality control is not a vibe and not a single accuracy number from a model you have not trained yet.
A person reads a spot-check sample: random rows plus the rows the rules found odd. The report names gaps instead of hiding them — a class with twenty examples, a source that arrived late, a field that is missing half the time. Gates are written: what must pass before I call the dataset ready. I do not “fix” a failed gate by deleting the test slice.
Acceptance is the report, the sample with reviewer notes, and a script that reprints the same counts from the delivered files. A report on a corpus that is already built is often 1–3 weeks. If the numbers say the data is not ready, the next work is cleaning or fine-tuning data, not a launch.
