One command, six layers, and a model allowed to judge only one thing
We run seven trading strategies at once across six infrastructure layers. Something is always slightly wrong, and almost none of it matters. Reading the logs to confirm that took forty minutes a day. This is the routine that replaced it: collection the model cannot vary, a judgement it is allowed to make, and a dated claim about the next run so the routine can be caught being wrong.
1. The problem, and who has it
We run multiple strategies at once, on purpose. They trade different venues, different chains, different time horizons. One of them stopping is not an incident. It is the expected state. Venues change mechanics, an RPC provider quietly degrades, a collector loses a websocket and reconnects into a loop, a data refresh silently does not run.
The problem was never collecting that. We had logs, metrics, canaries and a data lake. The problem was that reading them took forty minutes and the answer was almost always "fine". Any team running more than a couple of services with existing telemetry has the same forty minutes somewhere in its day.
2. How we built it
The whole command is check status pls. What comes back is a verdict on seven strategies across six layers. The collection is one shell routine with a fixed set of queries: managed containers, virtual machines, health canaries, collector error logs over the last thirty minutes, lake freshness, and the feed. The model does not choose what to look at. It cannot skip a layer because it seems fine, and it cannot produce a log line that was not returned.
What it is allowed to be smart about is only the last step. Is this error new or aging. Is this staleness expected or real. Does a human need to know. That is the entire judgement surface.
3. What we found
Here is an actual reply, unedited:
Status: healthy across all 6 layers
[1] Cloud Run: bloomer pristine (0 SELL_ERR), 72 RPC_FAILOVER (handled). The 1
CURVE_REAL_SETTLE_ERR is the same lone 08-20 17:34 blip aging out — benign.
[2] GCE: 2 dlmm-lp ERRORs = the aging rh-stocks-check stragglers (14:10 & 15:10 08-20,
pre-disable — none since, disable held; they drop off the window tonight).
[3] Canaries: radar / sushi / rh all green.
[4] VM collector errors (30m): clean — reconnect patches holding.
[5] Lake: kv_tokens + cache.parquet ~23h old — within the daily cadence (last ran 08-20
14:30; today's refresh due 14:30, ~45 min out). Not stale.
[6] Feed: ok — last item 6.8h ago, 15/24h.
Nothing needs action.
The useful part is not the six layers. It is the word "benign". Three errors appear in that report, and all three are correctly ignored, each for a different reason.
The one settlement error is a single event from two days ago, still inside the query window, already seen. The two virtual-machine errors are stragglers from a job we disabled the day before. They predate the fix, none have occurred since, which means the fix is holding. The lake being 23 hours old is not staleness, because the refresh runs daily and the next one is 45 minutes away.
None of those judgements can be made from the log lines alone. Each requires knowing what we changed and when. A dashboard has no memory of your intent. It will show you the same two red counts today that it showed yesterday, and it cannot tell you that yesterday's fix is the reason they stopped growing. That gap is where most of the forty minutes used to go.
| signal in the report | number | verdict and why |
|---|---|---|
| Sell errors on the lead strategy | 0 | Pristine; nothing to judge |
| RPC failovers | 72 | Handled by design; volume, not a fault |
| Settlement errors | 1 | Aging out; same blip from two days ago, already seen |
| VM errors | 2 | Stragglers from a disabled job; none since the fix |
| Lake age | 23 h | Within a daily cadence; next refresh 45 minutes out |
| Feed items in 24 h | 15 | Normal; last item 6.8 hours ago |
The report also commits to a prediction. It ends with what it will check next cycle: whether the 14:30 lake refresh actually lands, because that job has a history of not starting. A report that only describes the present cannot be caught being wrong. One that makes a dated claim about the future can be, and the next cycle either confirms it or does not.
4. What it is not
It does not catch anything that produces no signal at all. A strategy that runs, executes cleanly, logs nothing unusual and is quietly unprofitable passes all six layers. Reconciling intent against outcome is a separate routine, and the two do not substitute for each other. We learned that the expensive way, on a strategy whose logs were spotless for weeks.
It is also one shop's routine over one stack. The six layers are ours. The split between collection and judgement is the part that transfers.
5. What to do with this
- Write the collection down as a fixed script so it cannot vary between runs. Same checks, same order, every day.
- Let the model classify only what came back: new or aging, expected or real, needs a human or not.
- Give the routine the change history. "Disabled yesterday" is the fact that turns two red counts into a non-event.
- End every run with a dated claim about the next one. Grade it next cycle.
- Run a separate reconciliation of intent against outcome for anything that spends money. This routine will not see a spotless, unprofitable strategy. That one is described at agent-verification.
Reproduce it
You do not need our tooling. Take the six queries you already run when something feels off, put them in one script in a fixed order, and hand the output to a model with three questions and your last week of changes. If the value is not obvious after five days, the forty minutes were not the problem.
Sources: our own status routine and its reply on 20 August 2026, quoted verbatim above. Related: why correct agent code still does the wrong thing · fourteen ways AI agents fail with every check green · the atlas of machine memory.