Stays Up

Your AI coding agent is lying to you. Tests will not tell you.

If an agent writes code on your team, most of what gets merged is no longer read line by line. The code compiles, the tests pass, and the job still does the wrong thing. This is what we run instead of trusting the green tick: permission lists that grew one incident at a time, a reconciliation step that compares what the agent claimed with what actually happened, and the one line that replaced an incident.

1. The problem, and who has it

One of our CLAUDE.md files opens with a rule: never run the indexer without a 60-second timeout. It exists because an agent once ran it without one and left it running. The chain ingestor polls RPC continuously. It burned through credits overnight on a machine nobody was watching.

The code was fine. It compiled. It did exactly what it was asked to do. It would have passed every test we could have written for it. That is the shape of the problem for any team that lets an agent touch production: the failure is not incorrect code, it is correct code doing something nobody wanted.

It matters more now than a year ago for a simple reason. Generation went up. Review did not. On our team, and probably on yours, most of what gets merged is no longer read line by line.

2. How we measured it

We did not run a benchmark. We counted what we had built in self defence. Two workspaces, both with an agent that can run commands. We listed the explicit permission rules in each, checked how each rule got there, and went through the incidents that produced them.

Then we read the code of an open-source integration platform we were evaluating, specifically its job scheduler and its metrics module, to see whether a standard, well-maintained system would catch the same class of failure. We are not naming it. Its documentation warns about the behaviour we describe, so nothing here is a secret.

Finally we looked at the last months of our own trading agents, where every action is reconciled against the chain, and counted what the reconciliation caught that the test suite had not.

3. What we found

The permission lists are the first number. One workspace has 459 explicit rules, the other 754. Almost every one was added after we wanted the agent to do something and the rule was not there yet. Almost all were added after we wanted the agent to do something and it was not on the list. The lists are a written admission that correct code can still do the wrong thing.

Tests share the blind spot. A test checks that the code matches its description. When the same agent writes the code and the test, the two can agree by construction. The suite goes green without anyone having checked whether the description was right in the first place.

The cleanest example is not ours. In the integration platform, sync jobs are capped. A job that hits its execution limit is interrupted gracefully and resumes from a saved checkpoint. But a job that never advances its checkpoint restarts from zero every time. It never finishes. The system reports success on each run while guaranteeing that the job will never complete.

Its metrics enum has 135 types. Tasks created, started, succeeded, failed, expired, cancelled, retried, dropped. Queue depth. Handler duration. Not one of them says whether an execution was interrupted. A sync that will never complete produces the same telemetry as a healthy one.

what we countednumberwhat it means
Permission rules, workspace A459Each one a thing the agent may run, added after it was needed or misused
Permission rules, workspace B754Same pattern, larger surface
Metric types in the platform's enum135None distinguishes an interrupted run from a finished one
Lines that replaced an incident1timeout 60 in front of the indexer

The published literature agrees with our small numbers. In the fourteen cases we collected at agents-fail, the decisive error in coding-agent runs lands at a median of step 7, 82 percent of failed recoveries keep executing without progress, and 26 percent of failed runs end by announcing success. A green result late in a run is the least trustworthy thing an agent says.

Reconciling what the agent claimed against what actually happened has caught more problems in our systems than the test suite has. Tests check the code. Reconciliation checks the world.

4. What it is not

None of this is a test-suite replacement. Tests still catch the bugs they are written for. Permission lists still let through anything that is on the list, which is how the RPC bill happened: the agent had permission to run the indexer, the code was correct, and there was no reconciliation for credit spend. The rule went into the context file afterwards, which is where most of our rules come from.

The numbers above are two workspaces and one platform. They describe the shape of the problem, not its rate across the industry.

5. What to do with this

Reproduce it

Count your own permission rules and read how each one got there. Then pick the job in your system that reports success on every run and check whether its checkpoint has moved in the last week. If it has not, you have found the same failure we did.

Sources: our own workspace permission files and incident notes, August 2026; the integration platform's public scheduler and metrics code; case numbers from fourteen ways AI agents fail with every check green. Related: the status routine that judges only the last step · the atlas of machine memory.