Synthetic data fills holes that real exports do not cover: a rare defect, a refusal the support team almost never writes, a language you have in policy but not in tickets. I generate rows from rules or a model, then mark every one as synthetic. Synthetic data does not replace the real distribution and does not quietly become the whole training set.
A human reads a sample before the batch is accepted. If the generator copies a validation example, that row is junk, not a gift. I keep a ratio you can see: how many real rows, how many synthetic, which class they were meant to support. Personal data is not “anonymized” by asking a model to invent a lookalike customer.
Acceptance is the generator note, the synthetic flag on each row, the sample review, and a count by class. A contained batch is often 2–4 weeks. The split still has to keep synthetic twins out of validation — train, validation, and test.
