Stays Up

A free recency rule matched Jev at deciding what an AI agent can forget

Long coding-agent sessions run out of context, and something has to decide which old tool outputs to throw away. A popular plugin hands that decision to Jev, TypeSafe's new decision model. We ran it on 60 real agent sessions and checked every decision against what the agent actually needed later. Jev ranked better than chance, but no better than keeping the newest outputs, and the plugin as shipped kept none of the 1,729 outputs it judged.

1. The problem, and who has it

Anyone running coding agents on real work hits the same wall. A session reads files, runs tests and searches the codebase, and every one of those tool outputs stays in the context until the window fills up. At that point the tool compacts the conversation, usually by asking the model to summarise it, and a summary can quietly drop the one error message or file path the agent needed ten steps later. Nothing fails loudly. The agent just works from a worse picture of its own history.

So there is an obvious appeal in a different approach: keep everything verbatim, and only delete the old tool outputs that no longer matter. The hard part is deciding which ones those are. That is a per-item decision, keep or drop, and it is exactly the kind of question the new wave of decision models was built for. Jev, launched by TypeSafe in mid-September, answers typed yes/no questions with a calibrated probability in well under a second. Within days a Claude Code plugin, fast-jev-compaction, put it in charge of compaction and collected about seven thousand GitHub stars.

A decision like this only earns its place if it beats the free alternative. The free alternative here is recency: when space runs out, keep the newest outputs and drop the oldest. That was the question we set out to answer.

2. How we measured it

We took 60 public sessions of the OpenHands agent solving real GitHub issues, from the nebius/SWE-rebench-openhands-trajectories dataset, 27 of them resolved. Each session was cut at its midpoint. The first half went to the plugin, running the author's own code unmodified with Jev 1.13, and the second half became the answer key.

A tool output counts as needed if, after the cut, the agent wrote a distinctive token that before the cut appeared only in that output: a file path, a qualified function name, a test name, an error type. When the agent later runs pytest tests/test_path_utilities.py::PathUtilitiesTest, that class name came from a file it had viewed earlier, so that view was needed. We checked 16 of these labels by hand and 15 held up. Our first label, "the agent re-ran the same command", turned out to be wrong, because agents re-run tests after an edit to get new output, not because they lost the old one. We dropped it before comparing anything, and wrote every rule down before running the comparison.

The comparison itself holds the budget fixed. At 10%, 25% and 50% of the old output size, we kept outputs in the order Jev ranked them and, separately, newest first, and counted how many of the needed outputs survived. Across the 60 sessions that meant 1,729 old outputs, 135 of them needed later. The whole run cost about three cents of Jev, at a median of 0.65 seconds per compaction.

3. What we found

Budget keptJev rankingKeep newestDifference, 95% CI
10%0.200.14+0.06 [−0.02, +0.14]
25%0.460.37+0.09 [−0.03, +0.20]
50%0.660.69−0.03 [−0.14, +0.07]

Jev's judgement carries some signal. Its keep scores separated needed outputs from the rest better than chance, with a ranking score of 0.68 where 0.5 is a coin flip, and it tended to rank file views and edits above pwd and git log. But at every budget the confidence interval includes zero, so across 60 sessions we could not show that it keeps what the agent needs any better than recency does. After our first ten sessions Jev looked 12 to 21 points ahead, and that gap shrank to what the table shows as the sample grew, which is the usual story with early small samples.

The plugin's design never shows Jev the tool outputs themselves, only the calls and the conversation around them. We tested whether that was the problem by showing Jev the first 1,000 characters of every output, changing nothing else. The results did not move: 0.46 against 0.46 at the quarter budget, a ranking score of 0.63. Jev's scores seem to track the kind of call rather than what the output says.

The finding that matters most for anyone using the plugin is about the threshold. It keeps an output when Jev's probability is at least 0.5. Across all 1,729 outputs, Jev never went above 0.24, with a median of 0.14. So as shipped, the plugin keeps none of the old outputs it judges and deletes 75% of the characters in the session, including every one of the 135 outputs the agent needed later. At that point the decision is made by a fixed number that does not match the model's scale.

A decision model can rank sensibly and still decide nothing, when the threshold in front of it does not match the scale of its probabilities. Before trusting any gate, check what its scores actually look like on your own traffic.

4. What it is not

"Needed later" is a stand-in for what the agent needed. A dropped output can often be re-read, so a miss usually costs extra tokens and time rather than a failed task, and we did not run the behaviour check that would measure task success directly. The sessions come from one agent, OpenHands with a Qwen coding model, cut at one point, while the plugin was written for Claude Code sessions. And a result that does not clear its confidence interval is not proof that Jev and recency are equal, only that 60 sessions could not tell them apart.

This is one question, asked the way this plugin asks it, and it says little about Jev on other tasks. A lower threshold would keep far more, and the ranking does contain some information. What we can say is that the free rule is the bar, and on this question the model did not clear it.

5. What to do with this

Reproduce it

The sessions are public on Hugging Face (nebius/SWE-rebench-openhands-trajectories, rows 0 to 59) and the plugin is tamaratran/fast-jev-compaction on GitHub. The rules above are enough to rerun it: cut each session at its midpoint, label an output as needed when a token that appeared only in it shows up in the agent's later writing, and compare Jev's ranking with newest-first at the same character budget.

Sources: TypeSafe API and model documentation (Jev 1.13, $0.042 per million input tokens); fast-jev-compaction at commit e3f262a; SWE-rebench OpenHands trajectories. Related: AUC said the new model was better, volume fell 40% · fourteen ways agents fail with every check green · your coding agent is lying to you.