Contents
In brief
When a file is too big for a context window, an agent usually reads it in slices and hopes the answer sits in one of them, or searches a pre-cut index. Both lose text at the seams. The open-source TypeScript project Matryoshka leaves the whole document on the server — the book stays in the stacks — and hands back a short call slip instead of the pages. The Dev.to author wrote every command by hand, with no LLM API key and no local Ollama. Literal counts matched a separate check once the command rules were clear. The token-savings figures in the project README were not measured.
What happened
Matryoshka is by Dmitri Sotnikov (yogthos, author of the Luminus framework for Clojure). The README ties it to the Recursive Language Models paper by Alex L. Zhang, Tim Kraska, and Omar Khattab of MIT CSAIL. It is an independent implementation, not the paper authors' code. On 4 October 2026 the repository had 149 stars, 20 forks, and no open issues. The license is Apache-2.0. The latest npm release was matryoshka-rlm 0.2.40, dated 17 May 2026.
The model does not write JavaScript or Python. It writes Nucleus, a small S-expression language: (grep "ERROR"), a filter over the previous result, a count, a sum. The Lattice engine parses each command, type-checks it, and runs it against the document. Results stay in an in-memory SQLite store. The client receives a stub such as $grep_error: Array(1000) [preview...]. Real lines come back only through lattice_expand. The README claims "97%+ token savings" and "80%+ token savings compared to reading files directly" for the MCP server. Those are the project's numbers. The author did not benchmark them.
The lattice-mcp server needs no LLM at all, so the test covers the engine and the tool interface. It says nothing about whether a model would choose the commands well. The machine was Linux x86_64 with 8 vCPUs, about 15 GB of RAM, Node 20.19.2, and pnpm 10.33.4. The package was installed locally next to the official MCP SDK, not globally as the README suggests. Install finished in 2.5 seconds and used 220 MB of node_modules. pnpm 10 skipped native build scripts for the tree-sitter grammars. Left unbuilt, symbol listing still worked.
A short TypeScript client connected in 409 ms and reported lattice 0.2.40 with 12 tools: load, query, expand, close, status, bindings, reset, memo and memo delete, help, and two calls that prepare questions for a model.
Why it matters
Three inputs. War and Peace, Maude translation, from Project Gutenberg: 3.36 MB, 66,041 lines, 566,333 words, about 770,000 tokens (769,825 with tiktoken's o200k_base). That is past most context windows. An Apache access log from Elastic's examples: 10,000 lines, 2.37 MB, traffic from May 2015. Plus Matryoshka's own sources, for the code tools. Ground truth was computed first with grep and a short Python script.
Loading the novel took about 0.75–0.8 seconds. Later queries took between 1 and 430 ms.
| Question | What was asked | Matryoshka | Check |
|---|---|---|---|
| Chapter headings | search, then count | 365 | 365 |
| Book headings | same pattern, wide | 17 | 15 until the pattern was tightened |
| Mentions of Natásha | count matches | 1,213 | 1,213 (on 1,197 lines) |
| Pierre lines that also mention Natásha | filter | 44 | 44 |
| Where is "Genoa"? | search, then expand | lines 840, 1573, 9153 | the same |
| Opening quote | lines 840–846 | the seven-line "Well, Prince…" | exact |
| Mentions of Borodinó | count | 108 | 108 |
Grep is case-insensitive and counts matches, not lines. Two of the 17 "books" were ordinary sentences starting with a lowercase "book". The source runs grep with the gmi flags. A tighter pattern, ^BOOK [A-Z]+: , returned 15. After that, counts matched grep -oi.
Ranked search depends on spelling. (bm25 "battle of Borodino" 5) returned five lines that only said "battle". With the accent used in the translation, five lines all contained "battle of Borodinó". (fuzzy_search "Natasha Rostova" 3) found nothing; the accented form found her at once. The command the README calls semantic is TF-IDF cosine similarity, not embeddings. A query about being wounded at Austerlitz and looking at the sky missed the famous passage. The closest hit was "been looking at." on line 15862, seven lines earlier. Plain (grep "lofty sky") found the passage at line 15869.
On the log, 404 responses counted 213, Googlebot requests 543, and 200 responses 9,126. Summing the raw lines produced a meaningless 1,067,060.464, because sum takes the first number on each line — the start of the client address. After the byte field was extracted, the total matched: 2,735,455,845.
The first session against the novel made 27 tool calls. The tool responses added up to 5,625 characters. That is the author's own count, not a benchmark. The book stayed on the server. Only stubs and the lines that were asked for came back.
In practice
Asking a huge file this way is a reading-room request: you name the shelf, you get a slip, and only then do you ask for the volume. The slip is useless if you do not know the catalog rules.
Plain npx lattice-mcp from a folder where Matryoshka is not installed downloaded an unrelated npm package, lattice-mcp 1.6.2, which asked for LATTICE_API_URL and LATTICE_API_TOKEN. Use the binary that ships with matryoshka-rlm. The real server prints lattice-mcp v0.2.40.
The rest of the surprises were rules, not crashes:
- Spell words the way the document spells them, accents included. Lexical ranking does not fold them.
lattice_expandneeds a named handle.RESULTSworks inside a query and fails as an expand target ("Invalid handle: RESULTS").lattice_bindingslists names such as$grep_genoa. Patterns full of quotes get generic names, so the binding list is worth calling often.- Sum on raw log lines adds the first number. Pull the field out first, then sum.
- Loading
/etc/hostnamewas refused: the path is outside the working directory. A memo saved withlattice_memowhile the novel was open could still be expanded after switching to the log. - The test client does not support MCP sampling. A batch question over three 500-error lines returned one
[LLM_BATCH_REQUEST …]message with three ready-made prompts. The author typed["bot","bot","human"]and sent them back withlattice_llm_batch_respond. Filtering on "bot" counted 2. No model judged those lines. - On
lc-solver.ts(2,658 lines), listing functions returned 20, the same count as top-level function declarations, each with a line range and signature. The body ofevaluatecame back inline, 64,311 characters, and not as a handle, so a large function lands straight in context. - The call graph failed on 0.2.40 for all four source files tried: "No symbol graph available". Server logs showed
Graph.addEdgewith the same source and target. The graph library rejects self-loops, so a recursive function or a same-name call seems to stop the build. Four files do not say how widespread that is. - In
lattice-repl, searching for Austerlitz and counting gave 51, matching a separate check. Piping a load and a query together failed because the query arrived before the load finished. Passing the file as an argument worked. TherlmCLI, where a model writes Nucleus, defaults to Ollama on localhost, retries three times, and stops after about 9 seconds: three consecutive call failures, fetch failed. That is expected without a model.
Not tested: a real agent choosing Nucleus commands, rlm with a model, recursive queries, MCP sampling, compaction, resource limits, Claude Code wiring, HTTP and pipe adapters, program synthesis, extra ranking modes, multi-document loading, and any of the README's token-savings figures or the paper's benchmark results.
Takeaway
On literal tasks the engine was exact. Counts, lookups, and the opening quote matched the separate check once the semantics were clear. The document stays in the stacks. The desk gets a slip and only the lines you ask to fetch. The mistakes are in the catalog rules more than in the syntax: case-insensitive search and a sum of the first number will report 17 books and a nonsense byte total with full confidence. Ranked search is lexical. The call graph on 0.2.40 was the weak spot on the files that were tried; symbol listing still worked.
It is worth a try when an agent works over large logs, transcripts, or books and the answers need to be checkable, rather than pulled from pre-cut chunks. The author's next step is to put a real model in the loop and see whether it avoids these traps on its own.



Comments