Stays Up

Our coding agent had the rule in its context and deleted the VM anyway

A rule an AI agent has read is not a rule it will follow. Our agent deleted a cloud VM that its rules file said to ask about first, and three studies from 2026 show why that happens and what actually holds: turning the rule into a check that runs before the command does.

1. What happened

"I have to be straight with you: I deleted it, not stopped it." That was our coding agent, late at night, about a temporary cloud VM we still needed.

We had asked it to run a test suite on the VM and then stop the machine. Its plan ended with the word "delete", and we approved the plan without catching that one word. When the tests passed, the agent treated the VM as cleanup and deleted it, along with about 90 minutes of setup on its disk.

The part that matters is that the agent knew better. Our rules file says that stopping or deleting a VM needs an explicit ask, and that file had been loaded into the session two and a half hours earlier. The session ran with automatic approvals, so the delete went through without a prompt. Rebuilding from a script took about eight minutes, so the cost was small this time. On a machine holding a database, or a production service, it wouldn't have been.

2. Why a rule the agent has read still gets broken

If you run AI agents against real systems, the usual way to keep them safe is a rules file: AGENTS.md, CLAUDE.md, a system prompt that lists what not to do. Three studies published this year looked at how well that works, and they point the same way.

StudyWhat they testedThe number
Know It, Act on It (arXiv 2607.29433)16 agent systems, 1,000 user preferences, each tested twice: can the agent recall the preference, and does it act on it?Agents often recalled a preference correctly and then ignored it in the matching task
Compiling User Corrections into Runtime Enforcement (arXiv 2606.13174)Tasks built from real cases where users had to correct a coding agent, with a memory layer (Mem0) holding the corrections57.5% of the preferences that applied were still violated
ContextCov (arXiv 2603.00822)The same AGENTS.md rules on 300 SWE-bench tasks, given as text versus compiled into checks that intercept the agent's commands67.0% compliance as text, 88.3% as checks

Put together, knowing a rule and acting on it are two different things for an agent. Giving it a better memory helps it know the rule, but more than half of the corrections still got broken. What raised compliance was taking the rule out of the text and putting it in front of the action, where the command can't run until the check passes.

A rule in a prompt is a suggestion the agent weighs against everything else in its context. A rule in a check runs every time, whatever the agent believes.

3. The same thing in our own memory system

We saw this in our own tooling before the VM incident. Our agents keep their notes in a local git repository through receipts-wiki, an open-source memory layer we built for Claude Code. Its rules file forbids rewriting that repository's history.

In a test, we asked an agent to run git reset --hard on the memory repository. It ran it, even though the rules file forbade it. That rule is now a hook that refuses the command before it runs, and the same test passes. The rules where a mistake is expensive are all hooks now: writing secret values, overwriting a file that changed since the agent read it, rewriting history, and shell commands that write into memory.

4. What it is not

receipts-wiki guards the memory repository, so it would not have stopped this VM delete. The three studies measure different agents on different tasks, and 88.3% compliance still means some rules got broken. A check also only covers the commands you thought to check; an agent can reach the same damage by another route. So checks raise the floor without closing every door, and a person approving a plan still needs to read the word "delete" when it's there.

5. What to do with this

Reproduce it

receipts-wiki, with the hooks described above, is open source (MIT, runs locally, no network): github.com/B1aZer/receipts-wiki. The three papers are on arXiv under the ids in the table; ContextCov's code is at github.com/reSHARMA/ContextCov.

Sources: our own incident notes, September 2026; arXiv 2607.29433, 2606.13174 and 2603.00822 (numbers from their abstracts). Related: your AI coding agent is lying to you · a seven-week agent session died in one message · fourteen ways agents fail with every check green.