← All posts

Neural nets by recipe: compute the weights, don’t store every matrix

A Habr survey of hypernetworks, Kilobyte Models, and SeedLM — compact “recipes” instead of giant weight files, and what you pay for the savings.

Neural nets by recipe: compute the weights, don’t store every matrix
Contents

In brief

A long Habr survey looks at research that stores a short “recipe” instead of gigabyte-scale weight matrices — a rule that rebuilds the numbers at runtime. Think of a cooking card versus a freezer full of finished cakes: less space on the shelf, but you still have to cook before serving. The author walks from hypernetworks and coordinate generators to Kilobyte Models and SeedLM for LLMs, tracking three budgets along the way: disk, working memory, and compute cost.

What happened

The familiar path for a local model is simple: keep a weight file on disk, then load it into GPU memory — all at once or block by block. Qwen2.5-7B-Instruct in BF16 already lands near 15 GB; giants like Kimi-K3 push into trillions of parameters and terabyte-scale footprints even under a crude four-bit accounting. The survey asks a different question: can you run a network without storing every matrix element, if a compact description can rebuild those elements on demand?

The most direct answer is a hypernetwork: a small net that emits the weights of a larger one. On Wide ResNet for CIFAR-10, parameter counts drop from millions to fractions of a million, but classification error rises. Later work adds coordinate predictors such as NeRN and Big2Small: layer and channel indices go in, a convolution kernel comes out. On ResNet18 and ImageNet, accuracy barely moves (about 71.5% → 71.2%), and the on-disk package shrinks further when the recipe is paired with quantization — though one reported inference latency grew by roughly 30%.

Then the recipe gets shorter still. Kilobyte Models keep a PRNG seed plus a tiny trainable latent; a small CNN description fits in about two kilobytes with MNIST accuracy close to the full FP32 baseline. SeedLM applies a related idea to an already trained LLM: each weight block gets its own seed and mixture coefficients from a linear-feedback shift register basis. On Llama 2-7B, WikiText-2 perplexity worsens only slightly at 3–4 bits per weight; on Llama 3-70B the same recipe hurts more. Seed-Q and AWSRC add uneven bit budgets and a compact residual fix on top of an already compressed base — quality recovers somewhat, but none of this is a lossless “magic zipper” for an arbitrary checkpoint.

Why it matters

Disk compression and runtime memory savings are different promises. You can store a recipe in kilobytes and still materialize dense weights before a forward pass, using as much VRAM as before. Or you can generate a block, multiply, and discard — peak memory falls, but every restore adds compute. LLMs also carry a KV cache for prior tokens; that structure has its own footprint and does not vanish because the weights look clever on disk.

There is also a hard expressiveness ceiling. With a fixed decoder, a short description can pick only a tiny slice of all possible matrices. That is why Kilobyte Models train inside a family the generator can reproduce, instead of losslessly packing any pretrained checkpoint. For anyone hoping LLMs will follow Doom onto every toaster, the survey is a useful filter: some papers show FPGA wins from fewer memory reads, some report careful ResNet numbers, and almost none deliver “download a recipe, get ChatGPT in the browser.”

In practice

Treat the piece as a map of trade-offs — not a mandate to delete your safetensors. It helps if you pick weight formats, target edge hardware, or compare plain quantization with more exotic schemes.

  1. Track three budgets separately: file size, peak memory while serving, and the cost of restoring weights on each pass.
  2. Do not equate “fewer trainable parameters” with “fewer gigabytes on disk” — recipes still carry latents, indexes, scales, and leftover dense weights.
  3. For an already trained LLM, look at block-wise schemes like SeedLM: savings are local, and quality loss scales with model size.
  4. Tiny on-disk size without faster inference is a valid outcome; Kilobyte Models authors explicitly say they do not compress inference compute.
  5. Before promising a “two-kilobyte model,” check whether training happened inside that generator’s family or someone is trying to squash a foreign checkpoint after the fact.

Takeaway

Recipes instead of full matrices are already real research: hypernetworks, coordinate generators, seeds, and latents deliver genuine on-disk compactness and sometimes help hardware read less memory. Compactness almost always buys quality loss, restore-time cost, or a narrow model family. An LLM on every kettle is still far away — but the literature map on Habr is a clear place to start.

Comments

Loading comments…