Cleaning and normalization make a consolidated table safe to train on. I fix types, dates, encodings, unicode lookalikes, and the difference between an empty string and a missing value. Cleaning and normalization is not deduplication and not junk removal: a valid rare row stays, a broken encoding does not.
The schema is written in plain language: what each field means, which values are allowed, and what happens to rows that fail. I do not silently coerce a free-text note into a category. Time zones and currencies are named. Text is normalized only where the brief says case, whitespace, or punctuation must not change the label.
Acceptance is a schema, a before-and-after sample, and a count of rows dropped for a named rule. The script fails closed on a new unexpected value instead of guessing. A single table is often 1–2 weeks once collection exists. Duplicates are the next pass — deduplication.
