LLM dataset preparation builds the records a language model is meant to learn: an instruction, the context it may use, and the answer you are willing to defend. Datasets for LLM are not a dump of your wiki into a prompt. Each row has a source, a license note, and a reason it belongs in the task.
I check for contamination: public benchmark text and near-copies of the validation set do not sit quietly in train. I do not paste a famous eval set in to “raise the score”. If the model must answer from your documents at runtime, that product is RAG, not a bigger pretraining corpus. Supervised pairs for a later update are fine-tuning data.
Acceptance is a field contract (instruction, context, answer), a sample you can read, and a note on what was excluded as contaminated or unlicensed. A focused instruction set is often 3–5 weeks. The starting map is also in dataset engineering for model training.
