← All posts

JSONL, shards, and the path to the GPU: from one line to training

How large training sets are laid out: JSON vs JSONL, shards, compression, tokens, batches, and epochs — and why a terabyte of data does not need a terabyte of RAM.

JSONL, shards, and the path to the GPU: from one line to training
Contents

A truck drops one half-ton crate on the loading dock and someone says: “that’s your dataset — go train.” The forklift cannot lift it. You cannot open it safely. You cannot replace one broken box without unpacking the whole load. That is what a “400 GB single JSON file” feels like: the notebook dies on RAM, “train on everything at once” sounds impossible, and “how much memory do we need?” sounds like “how big a warehouse do we need for the entire port.”

Production ML pipelines work differently. Data is packed onto pallets — shards; each pallet holds boxes — one record per line in JSON Lines (JSONL); the conveyor to the GPU carries small batches, not whole pallets. The size of the dataset and the amount of data resident in RAM or VRAM at once are different numbers. This longread follows the path from one line on disk to one training step, and names the mix-ups teams usually make.

It continues the dataset engineering series with a focus on physical layout and loaders. Instruction-corpus design lives in the SFT longread; retrieval corpora are covered in RAG data preparation. Here: warehouses, pallets, and the conveyor.

Key takeaways

JSON is one document. JSONL is a stream of independent objects, one per line. Large sets almost always need the second.

A shard is a slice of the set, not a magic file format. A shard may be .jsonl, .jsonl.gz, Parquet, or another container.

Dataset size ≠ RAM. You can train on a terabyte by streaming shards and keeping only the current batch hot.

The model does not “read JSON.” A line becomes tokens; tokens become a batch; the batch hits the GPU. JSONL is storage and interchange, not the model’s language.

Token counts matter more than file gigabytes for training budget. One GB of JSONL yields different token volumes by language, fields, and tokenizer.

JSONL is a strong interchange and ingestion format. Columnar analytics and selective field reads often favor Parquet or formats such as WebDataset.

Train, validation, and test are separate directories with their own shards. A folder soup without a manifest breaks run comparison — see the series pillar.

One giant JSON is a crate you cannot move

JSON (JavaScript Object Notation) is convenient: humans can read it, every language parses it, APIs and configs speak it. A small document of objects, arrays, strings, and numbers is a fine choice.

Trouble starts when that “small document” is the only container for hundreds of gigabytes. Then it stops being a config file and becomes a monolith:

  • parsers often want the whole document or keep a huge array in memory;
  • streaming is awkward — there is no cheap “next record”;
  • parallelism fights one file and locking;
  • corruption in the middle can make the rest unreadable.

On the dock, that is the half-ton crate: familiar format, impossible logistics. Training and ETL need a different contract for how examples live on disk.

JSONL: one line, one box

JSON Lines (JSONL; also NDJSON, Newline Delimited JSON) is a simple contract: each line is a separate, self-contained JSON object. There is no wrapping array.

{"id":1,"text":"Hello"}
{"id":2,"text":"How are you?"}
{"id":3,"text":"Goodbye"}

Compare with a classic JSON array:

JSON:   [ {...}, {...}, {...} ]   ← one document
JSONL:  {...}\n{...}\n{...}\n     ← stream of independent records

For ML this matches “one example — one record”: a dialogue for supervised fine-tuning (SFT), a class label, a question–answer pair. A typical instruction record might look like this (field design is covered in the SFT corpus article):

{"messages":[{"role":"user","content":"What is Python?"},{"role":"assistant","content":"Python is a programming language..."}]}

You can read it as a stream without loading the whole set:

import json

with open("dataset.jsonl", encoding="utf-8") as f:
    for line in f:
        item = json.loads(line)
        process(item)

Why this helps in practice:

  • low peak memory on sequential reads;
  • easy append of new examples;
  • filtering without parsing a giant array;
  • natural splits into files for distributed workers.

This is not “JSON without brackets.” It is a different contract: record boundaries are line boundaries; objects are independent.

A dataset is a set of examples, not “a file”

A dataset for training is an agreed set of examples for a task — not “the file we downloaded.” One record may be text, a Q&A pair, an image path, a class label, or a message list. Volume easily reaches millions of rows and hundreds of gigabytes or terabytes.

While the team thinks “we have data.json,” they are thinking about a container. Once you have a contract for “what counts as one example” and “how we check a slice,” you have a data product — in the sense of the series pillar. Physical layout (JSONL and shards) serves that product: version it, copy it in pieces, feed it to a loader.

A shard is a pallet, not a format

A shard is a piece of a large set. When one dataset.jsonl grows to hundreds of gigabytes, you cut it:

dataset.jsonl          →   dataset/
  (500 GB)                 ├── shard-00000.jsonl
                           ├── shard-00001.jsonl
                           ├── shard-00002.jsonl
                           └── ...

All shards together are the dataset. Each line inside a shard is still one example. Record order and how you pack files can affect training if you do not shuffle on purpose.

The critical distinction: JSONL is a record format; a shard is a slicing strategy. One shard may be .jsonl, .jsonl.gz, Parquet, TFRecord, or a WebDataset archive. Confusing “we have shards” with “we have JSONL” is like confusing “pallet” with “cardboard box.”

There is no universal “right” shard size. You will see tens of MB, hundreds of MB, 1–10 GB. The trade-off depends on worker count, disk and network speed, object-store limits on object count, ease of reprocessing, and the cost of listing millions of tiny files. Size the shard for your pipeline, not a magic number from someone else’s repo.

Why large sets get sharded

Pallets exist for logistics, not for pretty filenames.

Memory. You need not keep the whole set in RAM: read the current shard or a window inside it, process, free the buffer.

Parallelism. Different processes or accelerators can read different shards at once:

worker 1 → shard-00000
worker 2 → shard-00001
worker 3 → shard-00002
worker 4 → shard-00003

Storage and transfer. Copy, backup, and re-upload in pieces. Re-clean only the dirty shards, not the whole terabyte.

Fault tolerance. Recreate a damaged shard; the rest stays usable.

Distributed training. Under data parallelism, workers often see different slices of the stream. Exact layout depends on the framework: “one shard = one GPU” is a special case, not a law.

Slicing data is not slicing the model. Dataset shards and model parallelism (layers on different devices) are different scaling axes.

Compression: wrapping the pallet

JSONL text compresses well: repeated field keys, similar shapes, boilerplate tokens. Shards are often stored as shard-00000.jsonl.gz: less disk and less egress from object storage. You pay CPU for decompression and get weaker random access than some binary columnar formats.

A common production trade-off: keep cold/warm shards compressed, and decide in the hot loader whether to cache decompressed chunks. If disk and network are the bottleneck, compression usually wins. If CPUs are already saturated unpacking tiny shards, enlarge files or change format.

From a line on the shelf to a step on the GPU

JSONL is the source representation. Training consumes numbers. A simplified path:

flowchart TB
  raw[Raw data] --> clean[Cleaning and filters]
  clean --> jsonl[JSONL]
  jsonl --> shards[Shards]
  shards --> loader[Data loader]
  loader --> shuffle[Shuffle]
  shuffle --> tok[Tokenizer]
  tok --> batch[Batches]
  batch --> gpu[GPU: forward / loss / backward]

A token is a text fragment (word, subword, symbol) mapped to an id by the tokenizer. The model does not read “Hello, world!” as a human does: it sees a sequence of ids. So one GB of JSONL ≠ a fixed token count: language, JSON boilerplate, answer length, and vocabulary change the budget. Plan cost and steps in tokens, not only folder size.

A batch is the group of examples (or truncated sequences) processed in one step. The dataset is all examples; the batch is what is on the GPU conveyor right now. Batch size and sequence length together pressure VRAM — roughly the product “how many sequences × how long.” One shard may hold thousands of records and many batches — shard ≠ batch.

An epoch is one full pass over the training set (all train shards in an agreed order or with shuffling). Reading all 100 shards is notionally one epoch; order details belong to the loader.

“1 TB on a machine with limited RAM” then looks like:

1 TB dataset
  → 1000 shards × ~1 GB
  → read a chunk of a shard
  → examples → tokens → batch on GPU
  → free buffer → next batch

The warehouse is still full. The conveyor only ever holds one party of boxes.

Shuffle: do not feed the conveyor one shelf at a time

If the set starts with a hundred thousand car texts, then medicine, then code, the model sees unnatural “seasons” within an epoch. Gradients shift in domain chunks. So you shuffle: shard order, buffered record shuffle, or more expensive full shuffles.

A perfect shuffle of a huge set is expensive in memory and I/O. In practice teams combine random shard order with a shuffle buffer in the stream. Shuffle quality is a resource trade-off, not a checkbox for perfection.

JSONL vs Parquet: which shelf for which job

Trait JSONL Parquet
Human readability high low
Debug simplicity high medium
Record streaming convenient possible
Compression good (esp. gzip) usually very strong
Columnar field reads no yes
Typical role interchange, ingestion, SFT exports analytics, lakes, selective columns

JSONL often stays the language of exchange: easy to open, fix one line, script, hand to an annotator. Parquet wins when you need three columns from a wide table on terabytes or dense typed packing. One pipeline can use both: labeled interchange in JSONL, feature marts in Parquet. “JSONL is always best for large sets” is as false as “Parquet is always easier for editing one record by hand.”

A catalog layout for your own project

Minimum warehouse discipline: separate raw, cleaned, and training splits:

dataset/
├── raw/
├── cleaned/
├── deduplicated/
├── train/
│   ├── shard-00000.jsonl.gz
│   ├── shard-00001.jsonl.gz
│   └── ...
├── validation/
│   └── shard-00000.jsonl.gz
└── test/
    └── ...

Do not train on raw. cleaned / deduplicated are reproducible stages. train / validation / test are different families for leakage control: eval shards must not “accidentally” enter training. Pin a version manifest (file hashes, schema, build date) to every run — otherwise metric comparison is theater; the frame is in dataset engineering.

A typical production contour:

raw → clean → dedup → filters → JSONL → shards → compression
  → object storage → data loader → shuffle → tokenizer → batches → GPU

For LLMs the path is the same: web and documents become raw, then JSONL and shards, then tokens. JSONL is not “how the model thinks”; shards are not a network layer. They are volume infrastructure. When volume and task justify fine-tuning at all — see when fine-tuning is needed.

A one-million-row worked example

Suppose 1,000,000 records at ~5 KB each → on the order of 5 GB before overhead. Layout:

100 shards × ~10,000 records
dataset/
├── shard-00000.jsonl
├── ...
└── shard-00099.jsonl

Each shard feeds the loader: records → tokens → batches → GPU. If the batch is 32 sequences, a 10,000-row shard yields hundreds of steps — not “one step per shard.” That is where the intuition “file = training portion” breaks.

Minimal Python tooling: stream read/write, line counts, split into N files, gzip, skip bad lines with a log. Even a small sharding script turns a monolith into pallets:

import json
from pathlib import Path

def shard_jsonl(src: Path, out_dir: Path, rows_per_shard: int = 10_000) -> None:
    out_dir.mkdir(parents=True, exist_ok=True)
    shard_idx, n_in_shard = 0, 0
    out = None
    try:
        with src.open(encoding="utf-8") as f:
            for line in f:
                if n_in_shard == 0:
                    if out:
                        out.close()
                    out = (out_dir / f"shard-{shard_idx:05d}.jsonl").open(
                        "w", encoding="utf-8"
                    )
                    shard_idx += 1
                out.write(line if line.endswith("\n") else line + "\n")
                n_in_shard = (n_in_shard + 1) % rows_per_shard
    finally:
        if out:
            out.close()

Production adds schema checks, hash manifests, and a ban on silently rewriting a published shard without a new version.

Common mistakes

  1. “JSONL is just JSON without brackets.” No: it is a stream of independent objects with line boundaries; parsers and contracts differ.
  2. “A shard is a special format.” No: it is a piece of the set; the container inside can vary.
  3. “1 TB dataset ⇒ 1 TB RAM.” No: you need streaming and budget for the batch plus buffers.
  4. “One shard = one batch.” No: a shard usually holds many batches.
  5. “One shard = one GPU.” Not necessarily: depends on loader and parallelism strategy.
  6. “JSONL is always best for large sets.” No: columns and analytics often prefer Parquet or domain containers.
  7. Train and test mixed in one folder without a manifest. Pretty scores and unreproducible leakage — see the series on data families.

What to try today

  1. Open your current “dataset.” If it is one huge JSON or CSV, estimate the cost of cutting into JSONL shards and write a target shard size for your worker count.
  2. Write a five-field manifest: row count, shard count, hash of the file list, one-record schema, build date.
  3. Split train/validation/test into directories even if validation is tiny: the habit is cheaper than leakage.
  4. Count tokens on a 1,000-line sample with your tokenizer — compare to gigabytes and stop planning training by folder size alone.

FAQ

Must we always use JSONL to train LLMs?

No. JSONL is convenient for exchange and many SFT exports. Pretraining and large corpora also use custom binaries, WebDataset, Parquet, and mixes. Choose by bottleneck: record debugging, columns, network, decompression CPU.

What shard size should we pick?

One that keeps object listings sane, processes in a reasonable time, and does not leave workers idle on a single huge file. Teams often start around hundreds of MB to a few GB and measure loader throughput.

Is sharding the same as data parallel?

No. Sharding is about storing and reading pieces of the set. Data parallel is about devices computing gradients on batch slices and synchronizing. They often appear together; they are different layers.

Why not judge training volume only by JSONL size on disk?

Because boilerplate fields, language, dialogue templates, and the tokenizer change token counts. Two 10 GB files can yield very different training budgets.

What is next in the series?

The dataset engineering overview and SFT corpus design. Natural follow-ons are field contracts and metadata schemas, then train/eval splits and leakage.

Next in the series

Physical layout without a field contract becomes “every shard has its own schema.” Next in the queue: record schemas, metadata, and contracts (dataset-schema-contracts-2026), then explicit splits and leakage control (dataset-splits-leakage-2026).

Comments

Loading comments…