One threshold change made every retrain worse for a month. Nothing failed.
We changed one number in production and did not touch the model. A month later, every retrain was worse than the last. No bugs, no pipeline errors, no alerts. This is what the number did to the training data, the experiment that isolated it, and the one distribution we now watch between retrains so it cannot happen quietly again.
1. The problem, and who has it
Most production ML systems have a filter upstream of the model. It decides which records are eligible: which customers get scored, which transactions are reviewed, which candidates a trading model is allowed to see. The filter serves two masters. It gates live decisions, and whatever it lets through becomes the training population the next time the model is retrained.
That second role is the one nobody watches. A filter change looks like an operational tweak. It is also a change to the training set, applied silently, with a delay of one retrain cycle. If the change is small or conditional, the training data can drift for weeks before anyone connects the two.
Anyone who retrains on data their own system selected has this exposure: recommendation, credit, fraud, ad ranking, and every trading system that retrains on its own eligible universe.
2. How we found it
The change itself was sensible. One eligibility threshold went from a fixed value to a dynamic one, tied to a market variable so it would stay meaningful across conditions. At the time of the change the dynamic value was stricter than the old fixed one. Nothing moved.
Then conditions shifted, the dynamic threshold became more permissive, and a new sub-population started flowing through. It was structurally different from anything the model had trained on. By the next retrain, roughly a quarter of the training samples came from that population.
We noticed retrains getting worse and blamed the market. We tried rolling windows, temporal decay, and excluding date ranges. None of it helped, which in hindsight was the clue: every remedy assumed the problem was in time, and it was in population.
The experiment that isolated it was simple. Retrain on the same date range, the same features and the same hyperparameters, but with the old fixed filter applied to the training data. Performance swung by more than 8 percentage points. One filter.
3. What we found
The obvious hypothesis was that the new data was bad. We checked, and it was not. Over the same period, the two populations had nearly identical win rates and mean returns that were statistically indistinguishable. The new population was noisier, with fatter tails in both directions, but its center was the same.
So including data that was not worse made the model worse. We evaluated both models on held-out data, split by population, to see where the damage was.
| what | number | meaning |
|---|---|---|
| Share of training set from the new population | about a quarter | Enough to reshape what the model spent its capacity on |
| Swing from the filter alone | 8+ percentage points | Same range, same features, same hyperparameters, old filter |
| Outcomes, new vs. original population | indistinguishable | Win rates nearly identical, mean returns statistically the same |
| Ranking on the original population | filtered model better | The model that never saw the new data ranked the old data better |
| Ranking on the new population | near random, both models | Training on it did not help predict it |
| Degradation before it was diagnosed | about a month | Every retrain in between was worse than the last |
The model that had never seen the new population ranked the original population better than the model trained on both. Meanwhile both models were near random on the new population. Training on it bought nothing there and cost something everywhere else.
The mechanism is interference. A model with finite capacity was asked to learn two overlapping distributions with different shapes. The splits it spent describing the noisy one were splits it could not spend on the clean one. No amount of temporal weighting fixes that, because the problem was never when the data arrived.
4. What it is not
It is not a data quality incident in the usual sense. The new records were valid, correctly labelled and, by outcome, no worse than the old ones. A data validation layer checking schemas and ranges would have passed every row.
It is not a drift alarm either. The feature distributions of the original population did not change. What changed was the mixture: who was in the training set, not what any individual record looked like. Standard drift detectors on individual features are not built to see that.
And it is not an argument against dynamic thresholds. The change was right for live decisions. It was wrong only in the role nobody had assigned to it, as the selector of the next training set.
5. What to do with this
- List every upstream filter that feeds the training set, and treat a change to any of them as a training-data change that needs a before-and-after comparison.
- Track the shape of the prediction distribution between retrains. If the spread, the tails or the mean shift while accuracy metrics look flat, something upstream changed the population.
- When retrains degrade, test population before time. Retrain on the same range with the previous filter first. Rolling windows and decay cost us weeks and answered nothing.
- Evaluate held-out data split by population, not only in aggregate. An aggregate score hides a model that got worse on the data you actually care about.
- Decide explicitly whether a new sub-population should enter training. Being eligible for a live decision is not the same as being informative to learn from.
Reproduce it
Take the training set behind your current model. Reapply the eligibility rule that was in force two or three retrains ago and count how many rows it removes. If the answer is not zero, retrain on the reduced set with everything else unchanged and compare on held-out data, split by whether each row would have passed the old rule. The difference between those two models is the cost of the filter change, and it is a number nobody in your pipeline is currently reporting.
Source: our own production incident and the filter-isolation retrain described above. Numbers are from that experiment. Related: the missing line that cost about $1K a day · your model is fine, your threshold is lying · fourteen ways AI agents fail with every check green.