Jev scores 86.6% on public questions and 36.7% on sealed ones
If you pick an AI decision model from a leaderboard, the number you see was earned on questions the builders could see too. JevBench is one of the few benchmarks that also asks questions it never publishes. We read its full table: across 91 models the median drop from public to sealed questions is 46 points, while a general-purpose LLM keeps 92.9% on the same sealed set.
1. The problem, and who has it
Decision models are the new, cheap layer many teams are putting in front of their products: a model that reads some state and answers a typed question, yes or no, pick one of these, score this, in well under a second. TypeSafe's Jev started the wave in September, and dozens of open alternatives followed within days. Anyone choosing between them will look at a leaderboard first, and the question that matters is how much of the headline score survives on the hard cases in your own product that no benchmark has seen.
2. Where the numbers come from
JevBench is run independently by Benchmark Heaven, with its harness, public tasks and scoring rules open under MIT. Version 1.4.2 asks 93 systems 534 public decisions and 308 sealed ones that are never published, and reports accuracy on both, along with calibration, speed and cost. We read the public results table as it stood on 24 September 2026 and computed the drop for the 91 systems that report both accuracies. We did not rerun anything; every number below is theirs, and the summary statistics are ours.
One detail changes how the drop should be read. The sealed questions come from the benchmark's harder families: long policy documents, trade-offs, deliberately ambiguous cases, traps, paraphrases and safety judgements. The public set also contains easier tiers. So part of any drop is difficulty, and part is whether a model generalises beyond what it has seen.
3. What we found
| System | Public (534) | Sealed (308) | Drop |
|---|---|---|---|
| Jev 1.13 (TypeSafe), ranked #2 | 86.6% | 36.7% | 49.9 pts |
| decider-4b v2, ranked #1 | 83.5% | 34.7% | 48.8 pts |
| Laya | 58.4% | 30.8% | 27.6 pts |
| CLM-8B | 40.7% | 24.0% | 16.7 pts |
| GPT-6 Luna (general LLM baseline) | 99.1% | 92.9% | 6.3 pts |
Across the 91 systems, the median drop is 46 points and the mean 39.6. Seventy-four of them lose more than 25 points, which is the threshold above which JevBench already penalises a model's intelligence score. Only seven lose 10 points or fewer.
The general model is what makes the comparison fair. GPT-6 Luna answers the same sealed questions at 92.9%, so those questions are answerable, and the collapse of the specialised models is not just the sealed set being impossible. The ranking combines accuracy with calibration, speed and cost at equal weight, which is why a system can lead it while answering about a third of the sealed questions correctly.
4. What it is not
A sealed set of 308 questions is small, and difficulty and generalisation are mixed together in every drop, so none of this is a precise measure of either. It is not a verdict on Jev or any other model for your task: some of these models may be exactly right for simple, high-volume decisions where speed and cost matter most. And these are one benchmark's numbers as of one date; the table changes as systems are resubmitted.
5. What to do with this
- Before trusting a decision model, run it on a held-out sample of your own past decisions with known outcomes, including the hard ones, and compare it with the simplest rule you could write.
- Read public and held-out results side by side wherever a benchmark reports both, and treat a gap above 25 points as a warning.
- Weigh speed and cost against accuracy on your hard cases, not on the average case.
We saw the same on a real task. We gave Jev 60 public coding-agent sessions from the SWE-rebench OpenHands set and asked, for each old tool output, whether the agent still needed it. Measured against what the agent actually used later, and at the same token budget, Jev's ranking kept 46% of the needed outputs and a plain keep-newest rule kept 37%; across 60 sessions that difference was not distinguishable from zero. A strong leaderboard position did not turn into an advantage over the simplest rule on work with real outcomes.
Reproduce it
The full table, with public accuracy, sealed accuracy and the gap for every system, is on the JevBench page; the harness and public tasks are in its open repository. The drop statistics above are a median and mean over that table's 91 rows with both scores.
Sources: JevBench v1.4.2, Benchmark Heaven, benchmarkheaven.com/jev-models (read 27 September 2026, table dated 24 September); harness at github.com/fstandhartinger/jevbench (MIT). Related: AUC said the new model was better, volume fell 40% · closing our trading desk · fourteen ways agents fail with every check green.