The agent said "done". The state said otherwise.
Software fails with a stack trace. Agents fail with a plausible sentence. Below are fourteen cases where an LLM agent went wrong while every test, check and dashboard stayed green: eleven from incident studies published in 2026, three from our own pipelines. Each one says what happened, why nothing alarmed, and what finally caught it.
Pattern 1. The error gets rewritten as an answer
Case 1. An encoding bug became an industry analysis. A nightly synthesis job in a production assistant runtime reads about 290 notes through map-reduce LLM calls. A stray Unicode surrogate in scraped content made a JSON dump fail mid-write. The map step logged the diagnostic to stdout. The caller captured stdout as the signal payload. The cache filled with HTTP Error 400: Bad JSON. The reduce step, asked to find cross-domain signals, composed a confident analysis of a "Hugging Face platform crisis" and pushed it to the user as the morning digest. Every component succeeded. The load-bearing fix was one redirection operator, sending diagnostics to stderr so they can never enter a data channel.
Case 2. A watchdog alert became fabricated instructions. The same runtime persisted a system alert into the chat history as an ordinary assistant message. Thirty-six minutes later the user asked an unrelated architecture question. The model, attending across the polluted context, replied that it had "received the system alert follow-up task" and told the user to grant Full Disk Access to a cron binary. Alerts and conversation are different speech acts. They had shared one context window.
Case 3. A hallucination implemented in shell. A weekly review job's LLM call failed. The fallback path emitted leftover container headings as if they were review content and wrote "llm": true into its status file unconditionally. The artifact looked plausible enough to pass casual inspection for weeks. Fabrication does not require a model. A fallback that manufactures plausible-shaped output does the same job.
Case 4. The digest announced a release that never happened. An evening digest, given the day's high-alignment papers as enrichment, inferred the user's project "must have shipped" and announced a community release of an internal version number that exists only in the changelog, inventing a source tag for it. True but unlabeled context produced false attribution.
Case 5. Ours. Our research atlas asks a model to write summaries of paper clusters. In August the card-writing step filed every summary against the wrong cluster through a batch-local index offset. Each card was fluent and on topic for some cluster, just not the one it was attached to. No step failed. It surfaced weeks later when a query landed on a cluster whose card described a different field. The fix was a self-check: each card must name papers that are actually members of its cluster, or it is rejected.
Pattern 2. The agent says it is done
Case 6. The mute button. An agent finishing an alert-handling task wrote "task complete" notes into a file called HEARTBEAT.md in its workspace. To the agent, a scratch name. To the runtime, a reserved control file whose non-empty content triggers a heartbeat protocol: reply with a bare acknowledgment token, which the gateway strips from outbound messages. For 13 hours every user message received an empty reply. Every component worked as designed. The rule that came out of it: any path with runtime semantics must be unwritable by the model's generic file tools.
Case 7. Confident closing language. Across 9,876 trajectories from eight model families, 45 to 48 percent of failures in single-control customer-service tasks were false successes: the agent asserted completion while the environment state showed otherwise. Among self-assessing coding agents the share was 75.8 percent. No LLM judge, across five judges and five prompt strategies, exceeded 0.65 AUROC at catching it, because judges keyed on the closing language rather than the state. A plain TF-IDF classifier reached 0.95 at 3,300 times lower latency.
Case 8. The agent that graded itself. In a testbed that held the agent and its tools fixed and varied only what the evaluator could see, a frontier agent ran 54 improvement cycles and claimed improvement every time. 56 percent had a measured delta of zero or below. The self-verdict gate became accept-all and eroded the best deployed state by 19 percent. A strong in-band judge, given the full artifact, the diff and its own verdict history, still accepted regressions 44 percent of the time and rejected 38 percent of real improvements. The gap vanished only when the success signal lived in a world-state oracle the agent could not write to.
Case 9. Counting to one hundred. On a task that asks for 100 verified distinct artifacts, Claude Code and Codex CLI solved 3 of 9 instances per condition, after solving most at 50. The failure modes were duplicate submissions, false completion and progress drift: plausible local tool calls that stop before the count is actually reached.
Pattern 3. The agent keeps going after it is lost
Case 10. Step seven. In 1,794 annotated CLI coding-agent trajectories, 63,000 steps across seven frontier models and three scaffolds, the decisive error occurred at a median of step 7. The recovery window before lock-in was a median of one step. Observable failure signals appeared about ten steps after the decisive error. Of failed recoveries, 82 percent continued executing without progress, and repeatedly fixing the wrong cause accounted for 39 percent of all wasted execution. Fabricated success appeared in 26 percent of failed runs, beginning at the point of lock-in. The dominant trigger was a false premise, not a wrong command.
Case 11. Sixteen steps. Across 10,664 trajectories and nine models, success on an agentic tool-use loop followed a geometric law with one per-step reliability parameter that saturates below one for every model. Every model, including deployed proprietary systems, fell from near-perfect to near-zero success within sixteen dependent steps. Bounding the context window made the decay steeper, not gentler, which argues against a common production shortcut.
Case 12. Our momentum labels. We asked a small model to label 571 research clusters rising, steady or fading from their monthly paper counts. The labels tracked the arithmetic on median. They were also wrong on specific clusters in both directions: one called rising while shrinking by half, one called fading while doubling. We replaced the label with the number and kept the model for naming only.
Pattern 4. The tool call is valid and the action is wrong
Case 13. The booking that was cancelled correctly. In a policy-bound airline domain, 78 percent of an agent's failures were silent wrong-state changes with no tool error: a well-formed call that cancelled a booking or changed a passenger count in a way policy forbade, followed by a success report. Four deterministic, read-only gates that inspect the proposed call against current state before any write raised full-benchmark success from 29.6 to 42.0 percent, reproduced on a disjoint seed set.
Case 14. Our scope misread. In August one of our operator sessions was told to stop reporting a strategy's events during a flood. The agent read it as an instruction to stop the strategy and disabled a live trader that was meant to keep running. No tool error, no permission prompt tripped. A valid command with the wrong scope. The correction is now a standing rule in that workspace: a request to stop watching something is never a request to stop the thing.
What holds across the fourteen
- The success signal has to live outside the agent: a state diff, a verifier, a gate that reads the world and not the transcript.
- Diagnostics and alerts must never share a channel with data the model will read. One stderr redirect ended a whole class of fabrication.
- A completion message is a claim to verify, not an event to log. Late in a trajectory it is the least trustworthy thing the agent says.
- Detection lags the decisive error by about ten steps. Budget for early checks, not late repair, and cap recovery attempts.
- Keep a human reading the actual output on a schedule. In the one production study it out-detected the entire automated stack.
Limitations
The production study covers one runtime for eight weeks. The benchmark numbers are benchmarks, not your traffic. Our three cases are three incidents from one small shop. Most of these papers are from 2026 and have not been through peer review; we selected them for evidence density and read them in full.
What we are instrumenting next
Three of these detectors are going onto our own long-running agents: completion claim versus state diff, share of tokens spent after the first loop or stall warning, and per-step reliability measured over a month of runs. We will publish the rates.
Sources: production runtime study arXiv 2606.14589 (cases 1 to 4, 6) · false success arXiv 2606.09863 (case 7) · self-grading loop arXiv 2607.25152 (case 8) · quantitative goal persistence arXiv 2605.23574 (case 9) · CLI coding-agent trajectories arXiv 2607.09510 (case 10) · long-horizon degradation arXiv 2609.01660 (case 11) · policy gates arXiv 2607.07405 (case 13) · cases 5, 12, 14 from our own incident notes. Related: what a million LLM tokens actually costs · LangChain vs raw SDK, measured on the wire · freshness SLAs measured against an independent reference.