10,000 trades, one strategy, retired in profit. The full life cycle, measured.
Most write-ups about automated trading show the good months. This one shows the whole arc: launch, compounding, plateau, give-back, and the decision to retire a strategy that was still in profit. It is about the instrumentation that made that decision possible, and it applies to any system whose edge depends on a market that can move away from it.
1. The problem, and who has it
A strategy that has worked for months starts working less. Every operator faces the same question at that point: is the model broken, is the execution broken, or is there simply less of the thing you knew how to trade? The three have different fixes and, without instrumentation, they look identical from the P&L curve.
This applies beyond trading. A fraud model, a recommender, a lead scorer: each has an edge that depends on a population that can change. The tooling that separates "our system degraded" from "the world moved" is the same in every case.
2. How we measured it
The system went live with real capital in early January, at deliberately small size. Every trade ran twice. Once as a simulated, zero-cost baseline, and once with real money. That gave us matched pairs: the same signal, the same moment, one paper outcome and one real outcome. Over the life of the strategy we collected more than 10,000 of them.
The difference between the paper result and the real result on each pair is the execution gap. We tracked it as a rolling average from day one. Alongside it we tracked trades per day, the distribution of per-trade P&L, and cumulative P&L. Those four series, read together, are the whole story.
3. What we found
The execution gap started noisy, compressed to around 10 percent per trade, and then stayed flat as we scaled the size up. That flatness is the single most important measurement in the project. It told us two things at once: the edge was real, because paper and real agreed once noise settled, and the size had room, because scaling did not widen the gap.
The per-trade P&L distribution looked the way this kind of market makes it look. We lost more often than we won. The winners paid for everything. No single trade mattered to the outcome. The system was the edge, not any position in it.
Cumulative P&L rose through January, February and March on thousands of small trades compounding at small size. Then the curve flattened. Then it gave some back. The give-back is the part we would show first to anyone evaluating a strategy, because it is where the instrumentation earned its keep.
Trades per day is what explained the give-back. At peak the strategy was taking more than 150 trades a day. By May it was taking a handful. Same model, same thresholds, same pipeline, same execution gap. The flow we were trading had dried up and the venue mechanics had shifted underneath us. There was less of the thing we knew how to trade, and the model was correctly declining to trade what was left.
| series | what it showed | what it meant |
|---|---|---|
| Execution gap | noisy, then about 10% a trade, flat as size grew | edge real, size had room |
| P&L distribution | more losers than winners, winners paid for all | no single trade mattered |
| Cumulative P&L | compounding, plateau, partial give-back | something changed in spring |
| Trades per day | 150+ at peak, a handful by May | the flow went away, not the model |
The honest conclusion from our own evaluation pipeline was that the edge had not degraded. It was gone, because its input was gone. So we retired the strategy instead of forcing it. It still runs in paper, still accumulates data, and if the regime returns it is ready to go back on the same day.
4. What it is not
One strategy, one venue, one year. The specific numbers, a 10 percent execution gap and 150 trades a day, belong to that market and will not transfer. The retirement call was made from our own evaluation pipeline and is a judgement, not a theorem. A different team with a different cost of capital might have kept it live longer.
We are also not claiming the strategy was optimal. It was profitable, measurable, and retired while still ahead, which is a different and rarer claim.
5. What to do with this
- Run every real decision alongside a zero-cost simulated twin from day one. The gap between them is the only clean measurement of execution quality you will ever get.
- Watch the gap as a rolling series, not a summary number. Flat as you scale means the size has room. Widening as you scale means you have found your capacity.
- Track volume of opportunities separately from quality of decisions. A model that trades less because there is less to trade is doing its job.
- Decide the retirement rule before launch: what the series must show for you to stop. It is much harder to write that rule while the curve is giving back.
- Do not hold one strategy. Hold a portfolio of uncorrelated ones and rotate capital from the fading to the growing. We call ours the Hydra: when one head stops producing, the others are still working and new ones are growing. Your strategy is not your edge. Your ability to replace it is.
Reproduce it
You do not need our market. You need a decision system, a simulated twin that sees the same inputs at the same moment, and a log that stores both outcomes per decision. From that table you can compute the four series above with a few lines of SQL: rolling execution gap, outcome distribution, cumulative result, and decisions per day. The retirement question answers itself once the four are on one screen.
Sources: our own trading system, real capital from early January 2026, more than 10,000 paper-real matched pairs, retired in the spring of 2026, when the flow it traded had gone and kept running in paper. Venue and successor strategies deliberately unnamed. Related: the calibration bug in the model behind this strategy · fourteen ways AI agents fail with every check green · the atlas of machine memory.