Stays Up

A missing line in the serving path cost about $1K a day. The model never raised an error.

Every ML system that trades, ranks or approves things has two pipelines: one that trains the model and one that serves it. They are supposed to preprocess data the same way. Ours did not, and nothing in the system said so. This is how we found the gap, what it was doing to our decisions, and the audit that catches this class of defect in any pipeline.

1. The problem, and who has it

A trained model is only as good as the inputs it sees in production. Those inputs pass through code that fills gaps, scales values and reorders columns before the model scores them. The training pipeline has its own copy of that code. The two copies are written at different times, often by different people, and they drift.

When they drift, the model keeps predicting. The system keeps acting on the predictions. Accuracy in the offline evaluation stays exactly where it was, because the offline evaluation runs the training pipeline. There is no exception, no warning and no metric that moves. The only symptom is that production makes worse decisions than the backtest promised, which is easy to blame on the market, on luck, or on the model being stale.

This applies to anyone running a model in production with a separate serving path: trading systems, fraud scoring, lead scoring, recommendation, credit decisions. If the preprocessing is not literally the same code in both places, the question is not whether there is skew but how much.

2. How we found it

We did not find it by looking for it. We had fixed an unrelated bug and wanted to confirm the fix had not changed anything else. So we ran an audit we had built for that purpose: take every prediction the production model made over a window, run the same model offline on the same stored inputs, and compare the two numbers trade by trade.

The audit is not sophisticated. It is a join on the trade identifier, a subtraction, and a sort by absolute difference. Its value is that it compares what production actually did against what it should have done, input by input, instead of comparing aggregate metrics.

Everything matched, except one trade. It was off by 0.002.

A difference of 0.002 on a probability is easy to wave away as floating point. We did not, because the audit had matched every other row exactly, so the model and the inputs were deterministic. Something had to be different about that one row. It had missing values.

3. What we found

The training pipeline filled missing values with zeros before handing the data to the imputer. The serving pipeline skipped that step. So in production, the imputer saw the missing values and did what an imputer does: it replaced them with the training medians. A feature that should have been zero was scored as the dataset average.

The one trade that surfaced in the audit was the one where the difference happened to be small. It pointed at the mechanism, and the mechanism was everywhere.

whatnumbermeaning
Rows flagged by the audit1The one that showed the mechanism; off by 0.002
Inputs affected in productionnearly halfAny row with at least one missing value was scored on medians the model never saw as inputs
Direction of the errorlow scored highInputs that should have scored low were lifted toward the middle and passed the threshold
Size of the fixone lineFill with zeros before the imputer, in the serving path
After the fixexact matchProduction predictions equal to offline predictions, row for row
Cost while it ranabout $1K a dayTrades entered that the correctly-fed model would have rejected

Nearly half of all inputs were being scored on values the model had never seen during training. In training, a missing value meant zero. In serving, a missing value meant "typical". For a feature where zero is the informative case, that is the difference between a rejection and an entry.

The direction mattered more than the magnitude. Because medians sit in the middle of the distribution, the substitution pulled low-scoring inputs upward. We were entering trades the model would have rejected had it seen the real inputs. The model was fine. What it was shown was wrong.

One line fixed it. After the fix, the audit matched every row exactly.

Train-serve skew produces no error and moves no offline metric. The only instrument that finds it is a row-by-row comparison of what production predicted against what the same model predicts offline on the same inputs.

4. What it is not

This is not a modelling problem. The model, the features and the threshold were all correct. Retraining would have changed nothing, because the training pipeline was the one doing it right.

It is also not something a test suite catches by default. Unit tests on the serving code would have passed; the code did what it was written to do. The defect lived in the difference between two pieces of code that were each correct on their own terms.

And the audit only works if production stores the inputs it scored, not just the scores. If your serving path logs the prediction and discards the feature vector, you cannot reproduce the prediction offline and this class of bug stays invisible.

5. What to do with this

Reproduce it

The audit needs three things: a table of production predictions keyed by decision identifier, the stored input vectors for the same identifiers, and the model artifact that was live. Load the model, score the stored inputs, join on the identifier, and print the rows where the difference is not zero. If that list is empty, your serving path matches your training path for that model version. If it is not, the largest difference is the place to start reading code.

Source: our own production incident, found through the backtest-versus-production audit described above. Numbers are from that audit. Related: your model is fine, your threshold is lying · how a config change silently degraded every retrain · fourteen ways AI agents fail with every check green.