Stays Up

AUC said the new model was better. Volume fell 40%. The threshold was lying.

If a classifier decides anything in your product, a trade, a fraud flag, a lead score, this will happen to you at the first retrain. The ranking metric holds, the backtest passes, and the system quietly starts doing less. We measured why on our own trading model, and the fix is two small changes to how models are evaluated and thresholded.

1. The problem, and who has it

We retrained our production classifier on fresh data. AUC held steady. The backtest looked fine. We deployed. Then trade volume dropped 40 percent. Same threshold, same pipeline, same market. The model had simply stopped passing candidates through.

The first instinct was to blame the market. That instinct is the expensive part, because it sends the investigation in the wrong direction for days. The market had not changed. The meaning of the model's numbers had.

Anyone running a classifier behind a fixed cutoff has this exposure. Tree ensembles produce a different score scale every time they are trained. If your pipeline says "act when the score is above 0.60", the retrain silently changes what 0.60 means, and no accuracy metric will tell you.

2. How we measured it

We put the old and the new model side by side on the same out-of-sample data and compared the raw score distributions, not the metrics. Then we asked a question the metrics never ask: when this model says 0.70, does the thing actually happen 70 percent of the time? We bucketed predictions by score and counted real outcomes per bucket, on out-of-sample data, for both models.

Finally we added a second metric to the evaluation pipeline, the Brier score, and compared what it said against what AUC said about the same pair of models.

3. What we found

The two models had the same mean score. Everything else about their distributions differed. The old model spread its scores with a standard deviation of 0.19 and went as high as 0.91. The new model had a standard deviation of 0.14 and topped out at 0.79. It had compressed everything toward the centre. It refused to make extreme predictions.

Our fixed threshold of 0.60 sat at the 85th percentile of the old model's distribution and at the 91st percentile of the new one. Same number, different selectivity. That alone explains the volume drop: the new model was passing a smaller slice of candidates through an unchanged gate.

measureold modelnew modelmeaning
Mean scoresamesamenothing visible at a glance
Std of scores0.190.14the new model compresses toward the centre
Max score0.910.79it never makes extreme calls
Where 0.60 sits85th percentile91st percentilesame threshold, stricter gate
Production volumebaselinedown 40%the visible symptom

Then the calibration check. Tokens the model scored between 0.60 and 0.70 had an actual hit rate of 21 percent. Tokens scored between 0.70 and 0.80 hit at 49 percent. Every bucket was overconfident, by 25 to 50 percentage points. A score of 0.65 meant roughly one in five, not two in three.

AUC does not care about any of this. It measures whether the model sorts good above bad. A model that outputs random numbers between 0.80 and 0.90 can have a perfect AUC if the ordering happens to be right. The probabilities can be nonsense and the metric stays green.

The Brier score is the mean squared error between predicted probability and actual outcome. It penalises overconfidence directly. On our pair of models it disagreed with AUC: AUC preferred the old model, Brier preferred the new one. The new model won on precision at matched selectivity. It got there by refusing to output scores it could not back up, which is exactly the behaviour that broke our fixed threshold.

There was a second finding on the way, from an earlier version of the same pipeline. The code that was supposed to calibrate scores computed a cumulative distribution and then applied its inverse. That is an identity function. The calibration step had been doing nothing since the day it was written, and every metric had been green throughout.

AUC tells you the model can rank. It does not tell you the probabilities mean anything, and it does not tell you that 0.60 still means what it meant last month. A fixed absolute threshold on a tree ensemble is a scheduled outage that fires at the next retrain.

4. What it is not

This is one production system and one pair of models. The specific numbers, 0.19 against 0.14, 21 percent in the 0.60 to 0.70 bucket, are ours and will not be yours. The shape is general: tree ensembles are not on a fixed scale, and most classifiers trained on imbalanced data are overconfident. The size of the drift is the part you have to measure yourself.

Brier is not a replacement for AUC. Ranking still matters when you act on the top slice. It is a second metric that catches the failure AUC is blind to, and the two should be read together.

5. What to do with this

Reproduce it

You need two model versions, one held-out window, and about an hour. Score both models on the window. Print mean, standard deviation and maximum of the scores. Compute where your production threshold sits as a percentile in each distribution. Bucket by score in steps of 0.10 and count outcomes per bucket. Compute AUC and Brier for both. If the two metrics disagree, or the threshold percentile has moved by more than a couple of points, you have reproduced our incident on your own system.

Sources: our own production classifier, two consecutive versions evaluated on the same out-of-sample window; the earlier calibration-code finding from the same pipeline. Related: the full life cycle of the strategy this model drove · fourteen ways AI agents fail with every check green · the atlas of machine memory.