JSONL, Parquet, and WebDataset are how a prepared corpus leaves my hands. I pick the format your trainer already reads, not the one that looks modern. JSONL fits row-wise text and chat records. Parquet fits columnar tables and filters. WebDataset fits images packed with labels into shards a loader can stream.
Each shard has a stable name, a schema or a sidecar that names the fields, and a small reader script that opens the first batch. I do not hide a one-off notebook as the only way in. Large corpora are sharded so a laptop can inspect a slice without downloading everything. The longer note on shards is in the blog: JSONL, shards, and the path to the GPU.
Acceptance is one file set that opens, a schema that matches the sample, and a documented command that reads the first shard. Format conversion of an already clean corpus is often 1–2 weeks. The records themselves still come from text datasets or image annotation.
