Everything we measured and published
We do not publish advice. Each page below is a measurement we ran on a system we operate, or a body of published work we read in full, with the numbers and the method on the page. Newest additions at the top of each group.
AI agents and reliability
- Jev scores 86.6% on public questions and 36.7% on sealed ones. Reading JevBench's table for 91 decision models: the median drop from public to never-published questions is 46 points, while a general LLM keeps 92.9% on the same sealed set. What a leaderboard can and cannot tell you.
- We ran 751 agent sessions against one shared memory and the writes do not collide. How many AI agents work the same project at once without overwriting each other: 356 notes, 24 changed a day, and a write gate that compares a content hash and refuses the save when the file moved since this session read it. Including the two gaps it does not close.
- A memory layer took one compaction's loss from half a session to a fifth. One real AI-agent compaction, 2,134 messages cut to 3% of their length: of 353 things dropped, git recovered 50% and the agents' own notes another 29%. The 22 that nothing recovered were all rejected options, the one thing neither the code nor the history can reconstruct.
- Our coding agent had the rule in its context and deleted the VM anyway. What three 2026 studies show about rules agents have read: memory alone left 57.5% of corrections violated, and the same AGENTS.md rules held 67.0% of the time as text and 88.3% as checks that run before the command.
- A free recency rule matched Jev at deciding what an AI agent can forget. Sixty real coding-agent sessions, 1,729 compaction decisions: Jev kept 46% of the outputs the agent needed later against 37% for keep-newest, a gap inside its confidence interval, and the plugin as shipped kept none because its 0.5 threshold sat above every score Jev gave.
- Fourteen ways AI agents fail with every check green. An agent that muted itself for 13 hours by writing to the wrong file, a self-grading loop that accepted everything, coding agents whose decisive error lands at step 7. Eleven published cases and three of ours, with what caught each one.
- Your AI coding agent is lying to you. Tests won't tell you.. Correct code that burned RPC credits overnight, 459 permission rules in one workspace, and a sync job that reports success while guaranteed never to finish. Reconcile claimed against actual.
- One command, six layers, and the word "benign". Seven strategies checked by a fixed routine where the model may judge only the last step: is this error new or aging. Forty minutes of log reading replaced, and a dated prediction attached to every run.
- Atlas of Machine Memory. 7,414 arXiv papers on memory and long context, 18 months, clustered into 248 topics with a card each. A third of the field is making the context window cheaper; nobody has solved the write path. Searchable down to the papers.
Models in production
- Your ML model is lying to you. AUC won't tell you.. A retrain with steady AUC cut production volume 40%. Scores between 0.60 and 0.70 hit 21% of the time. AUC measures ranking, not honesty; Brier and percentile thresholds do.
- One missing line cost about $1K a day. Training filled missing values with zeros; serving did not. Nearly half of all inputs were scored on values the model never saw. Found by comparing every production prediction to its offline twin.
- A config change that silently degraded every retrain. An upstream filter went dynamic, a new sub-population entered training, and performance moved 8 points from one filter. The new data was not worse, only noisier. Watch the shape of your predictions.
- After 100 SOL in profit and ten months, we closed our autonomous trading desk. From asking an LLM which memecoins to buy, to our own data lake and LightGBM, to switching it off: 84.5 SOL net over nearly 11,000 real trades, one trade worth a fifth of it, and the measurements that said stop.
- 10,000 trades. A strategy retired in profit.. Every trade run twice, paper and real, over 10,000 matched pairs. The execution gap that let us trust the numbers, the regime that ended it, and why we retired it instead of forcing it.
Cost, measured
- The Snowflake bill was 58x the work. Forty identical queries, a tenth of a second of execution, 1.35 credits on default settings and 0.023 after two changes. The diagnostic query, the fix and the mistake on the way, measured on a trial account.
- Cutting LLM inference cost: three levers. What a million tokens actually costs across providers and the order in which the levers pay. Numbers, not vendor slides.
- LangChain vs raw SDK, on the wire. Zero token overhead, and one default that turns tool errors into incidents. Measured request by request rather than argued.
Data, maps and regulation
- RAG on GitHub, mapped. Frameworks flat, agent memory tripled. The repos worth knowing, from a clustered map of the ecosystem rather than a listicle.
- What actually breaks when a data pipeline hits a billion events. Eight silent failures with the number that found each: 78% of lag on 1 partition of 64, 34% batch utilization at full saturation, 60% of writes dropped for 20 minutes behind a green health check. Plus the last hop rebuilt locally and broken four ways.
- "500 ms from the chain tip." Measured against what?. Freshness claims checked against an independent reference. How to know whether a data feed is as fresh as its SLA says.
- The EU AI Act was delayed. Your logging wasn't.. Which obligations already apply, which were deferred to December 2027, and why six-month retention moves the real deadline forward.
Want one of these run on your systems?
Fixed-price audits, one week each, the report and the detectors are yours. Tell us what you run and we reply in writing within one business day.
Tell us what you runAll pages are static and free to read. Related: what we do and what it costs.