Text dataset preparation turns documents, tickets, chats, and knowledge pages into records a model can learn from. I keep language, source, and a stable id on every row. Text datasets are not yet chat fine-tuning files: first the text has to be whole, encoded, and cut where a human would cut it.
I strip only the boilerplate the brief calls junk. I do not translate a corpus “to make it bigger” unless translation is the task. Segmentation follows the unit of the label: a ticket, a paragraph, a dialogue turn. Mixed languages stay marked instead of being forced into one. Personal data is masked when the brief says the model must not see it.
Acceptance is a sample of records with source ids, a note on how text was cut, and a count of rows dropped as empty or unreadable. Delivery format comes next — JSONL, Parquet, and WebDataset. A single corpus is often 2–4 weeks after the exports exist.
