Stays Up

One compaction cut 2,134 messages to 3%, and 79% of what it dropped was still recoverable

When an AI assistant runs long, the tool compresses the session into a summary and you are left guessing what went with it. We measured one real compaction against everything our agents had written down. Git held half of what disappeared, the notes held another 29%, and what nothing held turned out to be a specific kind of thing.

1. The problem, and who has it

Anyone running a long AI session hits this. The context window fills, the tool compacts, and a summary replaces thousands of messages. The work continues and feels fine. Nobody can say what was lost, because the thing that would tell you is the thing that was just replaced.

The usual advice treats this as a tuning problem: compact earlier, compact manually, write a better summary prompt. That assumes the summary is where the answer lives. We wanted to know something different. If a session is going to be compressed to a fraction of itself, how much of what it knew is still reachable somewhere else, and what exactly is the part that is gone for good?

2. How we measured it

We could put a number on it because our agents already write what they learn into a git repository while they work, one fact per file. That gives three artifacts for the same moment in time: the conversation before compaction, the summary that replaced it, and the notes as they stood that minute.

The subject was a real auto-compaction, not a staged one: session cc4459f4, compacted at 2026-09-30T14:31:17Z. It took 577,714 characters across 2,134 messages down to a 17,272-character summary. That is 3.0% of the original length.

The method was to extract every named thing the conversation had mentioned, which gave 353 of them: file paths, identifiers, flags, commit hashes, concepts, proposed field names. For each one that did not survive into the summary, we checked whether a later session could still reach it, either in the git history or in the agents' notes. Anything found in neither was hand-classified rather than counted automatically, because that residue is the whole point of the exercise.

3. What we found

where it survivedshare of what was droppedwhat kind of thing
Git history50%code, paths, commits. Free, and every team already has it
The agents' notes29%decisions, findings, reasons. Only exists if something writes it down at the time
Neither20%the residue, examined below

Four fifths of what the compaction dropped had not actually gone anywhere. The half held in git is the unremarkable half, and it is worth saying plainly that it is free: if your work is in a repository, the compactor cannot take your code or your commit history with it.

The 29% is the part that costs something to have. It is not code, so git never had it: the decisions, the findings, the reasons a thing is the way it is. That share exists only because an agent wrote it down while it was working, and it is the difference between losing half a session and losing a fifth.

4. The 20% that nothing held

Of the roughly 72 items in the residue, 14 were normalisation artefacts, the same string counted twice under different spellings. That left 58 real ones, hand-classified:

Those 22 break down as 9 frontmatter fields proposed and then dropped, 9 lint rules and concepts coined mid-design, and 4 topic names weighed during a migration. Every one of them is reasoning rather than fact. Nobody had decided them yet, so no note owned them, and they were not code, so git never saw them.

The compaction did not lose our facts. It lost the arguing. What vanishes is the option somebody raised and the group rejected, which is exactly the thing you need later to avoid re-arguing it.

That changed what we built. The fix was not a better summary and not a compaction hook. It was giving deliberation somewhere to live: a note type for a decision, which has to name the alternatives and cite the evidence, checked by lint so the convention is enforced rather than remembered.

5. What it is not

Three things here cut against us and are worth stating.

Our first number was wrong, by a lot. The initial reading of this same compaction put the loss at 85%. The real figure is 20%. The first pass had counted anything absent from the summary as lost without checking git or the notes at all. Nothing in our system caught that; running the check properly did. We would have published 85% if the measurement had not been redone.

Our diagnosis was also wrong on first telling. We initially claimed that a rejected option had no home in our memory at all. Challenged on it, we counted: 94 notes already recorded one, inside the note that owned the work. The real gap was narrower than the story we had reached for, and it is the narrow version above that the 22 items actually support.

This is one compaction, not a study. One session, one codebase, one team's working habits. The 50/29/20 split is what happened here. The shape of the finding, that facts survive and deliberation does not, is the part we would expect to generalise, and it is also the part we have the least evidence for.

We also tested and rejected three tempting fixes. Steering the summariser does not work: PreCompact and PostCompact cannot inject context, and additionalContext fails schema validation on both. Blocking auto-compaction does work, but the hook payload carries no token count, so a hook cannot tell whether there is headroom to continue into, and blocking blind near the limit trades a lossy summary for a dead session. Building a recorder to capture every future summary was the most appealing option and still the wrong first move, because it measures a problem that is already 79% handled.

6. What to do with this

Reproduce it

The memory layer is an open-source Claude Code plugin, receipts-wiki: notes are plain files in a local git repository, every change is a commit naming the session, and past conversations stay out of the live context until asked for. The measurement needs three things from the same minute: the pre-compaction transcript, the summary that replaced it, and the state of the notes. Extract the named things from the transcript, subtract those present in the summary, then look up each remainder in the git history and the notes. Classify the remainder by hand. The automated count is what produced our wrong 85%.

Sources: AdaMem, arXiv 2606.21144, on under-specified write policies in memory systems. LLM-generated design rationale precision, arXiv 2504.20781. Agent Zero Memory and the citation lock, arXiv 2608.29606. Related: running many agents against one shared memory · fourteen ways agents fail with every check green · the atlas of machine memory.