Our coding agent had the rule in its context and deleted the VM anyway
A rule an AI agent has read is not a rule it will follow. Our agent deleted a cloud VM that its rules file said to ask about first, and three studies from 2026 show why that happens and what actually holds: turning the rule into a check that runs before the command does.
1. What happened
"I have to be straight with you: I deleted it, not stopped it." That was our coding agent, late at night, about a temporary cloud VM we still needed.
We had asked it to run a test suite on the VM and then stop the machine. Its plan ended with the word "delete", and we approved the plan without catching that one word. When the tests passed, the agent treated the VM as cleanup and deleted it, along with about 90 minutes of setup on its disk.
The part that matters is that the agent knew better. Our rules file says that stopping or deleting a VM needs an explicit ask, and that file had been loaded into the session two and a half hours earlier. The session ran with automatic approvals, so the delete went through without a prompt. Rebuilding from a script took about eight minutes, so the cost was small this time. On a machine holding a database, or a production service, it wouldn't have been.
2. Why a rule the agent has read still gets broken
If you run AI agents against real systems, the usual way to keep them safe is a rules file: AGENTS.md, CLAUDE.md, a system prompt that lists what not to do. Three studies published this year looked at how well that works, and they point the same way.
| Study | What they tested | The number |
|---|---|---|
| Know It, Act on It (arXiv 2607.29433) | 16 agent systems, 1,000 user preferences, each tested twice: can the agent recall the preference, and does it act on it? | Agents often recalled a preference correctly and then ignored it in the matching task |
| Compiling User Corrections into Runtime Enforcement (arXiv 2606.13174) | Tasks built from real cases where users had to correct a coding agent, with a memory layer (Mem0) holding the corrections | 57.5% of the preferences that applied were still violated |
| ContextCov (arXiv 2603.00822) | The same AGENTS.md rules on 300 SWE-bench tasks, given as text versus compiled into checks that intercept the agent's commands | 67.0% compliance as text, 88.3% as checks |
Put together, knowing a rule and acting on it are two different things for an agent. Giving it a better memory helps it know the rule, but more than half of the corrections still got broken. What raised compliance was taking the rule out of the text and putting it in front of the action, where the command can't run until the check passes.
3. The same thing in our own memory system
We saw this in our own tooling before the VM incident. Our agents keep their notes in a local git repository through receipts-wiki, an open-source memory layer we built for Claude Code. Its rules file forbids rewriting that repository's history.
In a test, we asked an agent to run git reset --hard on the memory repository. It ran it, even though the rules file forbade it. That rule is now a hook that refuses the command before it runs, and the same test passes. The rules where a mistake is expensive are all hooks now: writing secret values, overwriting a file that changed since the agent read it, rewriting history, and shell commands that write into memory.
4. What it is not
receipts-wiki guards the memory repository, so it would not have stopped this VM delete. The three studies measure different agents on different tasks, and 88.3% compliance still means some rules got broken. A check also only covers the commands you thought to check; an agent can reach the same damage by another route. So checks raise the floor without closing every door, and a person approving a plan still needs to read the word "delete" when it's there.
5. What to do with this
- List the commands in your agents' reach that destroy something: deleting cloud resources, dropping tables, force-pushing, rewriting history, removing buckets.
- For each one, check where the rule lives. If it's only a sentence in a rules file, the studies above say it holds roughly two times in three.
- Move those rules into checks that run before the command: an approval rule in the agent's own settings that forces a question, a hook that refuses the command, or a wrapper around the CLI. The ContextCov result is what that move bought on 300 tasks: from 67.0% to 88.3%.
- Keep the rules file for everything else. It's still the best place for conventions and taste; it's just not a safety mechanism.
Reproduce it
receipts-wiki, with the hooks described above, is open source (MIT, runs locally, no network): github.com/B1aZer/receipts-wiki. The three papers are on arXiv under the ids in the table; ContextCov's code is at github.com/reSHARMA/ContextCov.
Sources: our own incident notes, September 2026; arXiv 2607.29433, 2606.13174 and 2603.00822 (numbers from their abstracts). Related: your AI coding agent is lying to you · a seven-week agent session died in one message · fourteen ways agents fail with every check green.