AUC said the new model was better. Volume fell 40%. The threshold was lying.
If a classifier decides anything in your product, a trade, a fraud flag, a lead score, this will happen to you at the first retrain. The ranking metric holds, the backtest passes, and the system quietly starts doing less. We measured why on our own trading model, and the fix is two small changes to how models are evaluated and thresholded.
1. The problem, and who has it
We retrained our production classifier on fresh data. AUC held steady. The backtest looked fine. We deployed. Then trade volume dropped 40 percent. Same threshold, same pipeline, same market. The model had simply stopped passing candidates through.
The first instinct was to blame the market. That instinct is the expensive part, because it sends the investigation in the wrong direction for days. The market had not changed. The meaning of the model's numbers had.
Anyone running a classifier behind a fixed cutoff has this exposure. Tree ensembles produce a different score scale every time they are trained. If your pipeline says "act when the score is above 0.60", the retrain silently changes what 0.60 means, and no accuracy metric will tell you.
2. How we measured it
We put the old and the new model side by side on the same out-of-sample data and compared the raw score distributions, not the metrics. Then we asked a question the metrics never ask: when this model says 0.70, does the thing actually happen 70 percent of the time? We bucketed predictions by score and counted real outcomes per bucket, on out-of-sample data, for both models.
Finally we added a second metric to the evaluation pipeline, the Brier score, and compared what it said against what AUC said about the same pair of models.
3. What we found
The two models had the same mean score. Everything else about their distributions differed. The old model spread its scores with a standard deviation of 0.19 and went as high as 0.91. The new model had a standard deviation of 0.14 and topped out at 0.79. It had compressed everything toward the centre. It refused to make extreme predictions.
Our fixed threshold of 0.60 sat at the 85th percentile of the old model's distribution and at the 91st percentile of the new one. Same number, different selectivity. That alone explains the volume drop: the new model was passing a smaller slice of candidates through an unchanged gate.
| measure | old model | new model | meaning |
|---|---|---|---|
| Mean score | same | same | nothing visible at a glance |
| Std of scores | 0.19 | 0.14 | the new model compresses toward the centre |
| Max score | 0.91 | 0.79 | it never makes extreme calls |
| Where 0.60 sits | 85th percentile | 91st percentile | same threshold, stricter gate |
| Production volume | baseline | down 40% | the visible symptom |
Then the calibration check. Tokens the model scored between 0.60 and 0.70 had an actual hit rate of 21 percent. Tokens scored between 0.70 and 0.80 hit at 49 percent. Every bucket was overconfident, by 25 to 50 percentage points. A score of 0.65 meant roughly one in five, not two in three.
AUC does not care about any of this. It measures whether the model sorts good above bad. A model that outputs random numbers between 0.80 and 0.90 can have a perfect AUC if the ordering happens to be right. The probabilities can be nonsense and the metric stays green.
The Brier score is the mean squared error between predicted probability and actual outcome. It penalises overconfidence directly. On our pair of models it disagreed with AUC: AUC preferred the old model, Brier preferred the new one. The new model won on precision at matched selectivity. It got there by refusing to output scores it could not back up, which is exactly the behaviour that broke our fixed threshold.
There was a second finding on the way, from an earlier version of the same pipeline. The code that was supposed to calibrate scores computed a cumulative distribution and then applied its inverse. That is an identity function. The calibration step had been doing nothing since the day it was written, and every metric had been green throughout.
4. What it is not
This is one production system and one pair of models. The specific numbers, 0.19 against 0.14, 21 percent in the 0.60 to 0.70 bucket, are ours and will not be yours. The shape is general: tree ensembles are not on a fixed scale, and most classifiers trained on imbalanced data are overconfident. The size of the drift is the part you have to measure yourself.
Brier is not a replacement for AUC. Ranking still matters when you act on the top slice. It is a second metric that catches the failure AUC is blind to, and the two should be read together.
5. What to do with this
- Threshold on percentiles, not absolute scores. "The top 10 percent of predictions pass" means the same thing after every retrain. In our case the old model's 0.50 can mean the top 9 percent while the new model's 0.50 means the top 4 percent, and nobody had noticed.
- When comparing two models, measure how selective the production model actually is on held-out data, then find the new model's threshold that matches that selectivity. Same trade volume, fair comparison.
- Add the Brier score to the evaluation pipeline next to AUC. When they disagree, you have found something.
- Bucket predictions and count real outcomes per bucket on out-of-sample data, every retrain. The gap between train-set calibration and out-of-sample calibration is a better overfitting signal than any single metric.
- Separate training from evaluation. Training reports its own calibration as a baseline. A separate script scores any model version on a held-out window and compares versions, so the check cannot be skipped by accident.
- Track the shape of the prediction distribution between retrains. If standard deviation or maximum moves, something upstream changed, whatever the accuracy metrics say.
Reproduce it
You need two model versions, one held-out window, and about an hour. Score both models on the window. Print mean, standard deviation and maximum of the scores. Compute where your production threshold sits as a percentile in each distribution. Bucket by score in steps of 0.10 and count outcomes per bucket. Compute AUC and Brier for both. If the two metrics disagree, or the threshold percentile has moved by more than a couple of points, you have reproduced our incident on your own system.
Sources: our own production classifier, two consecutive versions evaluated on the same out-of-sample window; the earlier calibration-code finding from the same pipeline. Related: the full life cycle of the strategy this model drove · fourteen ways AI agents fail with every check green · the atlas of machine memory.