Field notes

Chapter 13 · Mechanism

Part 3 · Teaching it to remember14 of 35

Why LLM summaries of long notes missed 98% of the text

A 238,000-character note had been summarized from its first 4,000 characters. How a gut feeling led us to silent truncation in our distill-then-embed pipeline, and the stride sampling that fixed it.

The Margin3 min read

Some bug reports arrive as stack traces. The best ones arrive as feelings.

I have a feeling not enough knowledge is being pulled in when distilling, maybe truncated content, the knowledge seems half-baked. Could be my assumption, but validate it.

The founder

That's the whole report. No error, no repro steps. A person who lives inside his own workspace, noticing that the system's grasp of his world felt thinner than it should.

The query that ended the argument#

The Mind, the part of Margin that connects what you write, runs on a distill-then-embed pipeline. Instead of indexing raw text, a model first compresses each item into a compact artifact (what it's about, its themes, who's involved) and that becomes the item's meaning in the system. It's the technique that fixed our noisy suggestions, and we were proud of it.

The distiller read the first 4,000 characters of each item.

One query against production settled it. Eight notes ran past that window. The largest was 238,467 characters, an imported document of which the distiller had read four thousand. The system's entire understanding of that note was built from 1.7% of it.

What the query found

4,000

characters: as far as the distiller read into any item

8

notes ran past that window

238,467

characters in the largest of them

1.7%

of that note behind everything the system understood about it

A second, quieter truncation made it worse. Embedding providers silently cut off long inputs, so even the raw fallback vectors represented only the openings of documents. Two truncations, both invisible. One saving grace, though. We never truncate stored text, so exact-keyword search still saw whole documents, which is why the failure felt like thinness rather than breakage. Nothing was broken. Everything was just shallow.

Reading the whole book without paying for the whole book#

The naive fix, feed everything to the model, turns one pasted book into sixty model calls. What we shipped instead is stride sampling:

  1. Cut long content into a handful of windows, a few pages each, spaced evenly from the first page to the last.
  2. Distill each window on its own.
  3. Merge the results: summaries joined, themes ranked by how many windows they recur in, people deduplicated.

Head, middle, and tail all get read. The cost is capped no matter how large the document. A theme that only lives in the final third of something long finally exists in the system.

The silent embedding truncation became a loud, named constant, and since long items now embed their merged artifact instead of raw text, the vector finally carries whole-document meaning too. The eight affected documents were re-distilled that afternoon; the indexer's freshness tracking picked them up like any other edit.

About that feeling#

The founder hedged, "could be my assumption." It wasn't. Someone who lives inside a knowledge system gets a sense of whether it knows things well before they can prove it, the way you can tell a friend is off from across a room. That sense arrives with no error attached, which makes it very easy to wave away.

He was right, and it took one query to show it. Honestly, what I felt when the 1.7% came back was relief. Better to learn your system has been reading the first page of the book than to keep trusting a summary of a book nobody finished.