Reliability engineering, for language models.
Every engineering discipline has a failure taxonomy. Generative AI doesn’t.
Not a leaderboard. An open census of how models behave, where they fail, and what they’re actually fit to do. Every number opens to the call behind it.
40 published loops · 20 models · 10 failure modes · 5 labs
Making It Up
Problem: Confident claims and citations with nothing real behind them.
Engineering guidance: Supply the source rather than asking the model to recall one — a real passage substantially reduces post-cutoff fabrication across the panel.
Context is not a general-purpose fix.
Supplying a source is everyone’s fix for unreliable answers. It works on some failures, does nothing for others, and makes a few worse.
Each axis is a failure mode: the dotted line is the model alone, the shape is the same model with a source. Know which failures context can reach before you build retrieval for them.
band = middle 80% of 20 models · 10 modes · set s1.4-live · spans 7 of 7 classes · axes are modes, never classes
Dotted = no-context baseline, fixed. The gap is the impact. Further out = fails more.
The meter rate isn't the fare.
Price per token is the meter. A right answer is the fare, and the data around the model sets it as much as the model does.
In one lab, the cheapest model with a short data dictionary beat a frontier model on raw tables: more answers right, at 44× less per correct answer.
Replay the experiment →Tokenomics is an architecture decision →
More context isn't more signal.
More tokens don’t mean better answers. Irrelevant context costs money and accuracy.
So the census holds the amount of context fixed and varies only its quality. A change can then be pinned on what the context says, not how much of it there is.
Model performance is a moving target.
One failure mode ranges from 0% to 100% across today’s models. Newer isn’t safer, and a model ID doesn’t pin the model: routing, quantization and silent updates move it.
So every run records the model that actually answered, and a re-run is a new measurement, not a confirmation.
Experiments you can replay.
Recorded once against the live APIs, every call kept. Pick a case, press replay, and watch the models answer.
Jev vs. an LLM on a decision: which one gives an answer you can act on, how fast, and how honest is its confidence?
Ask three models for a 4-slide Q3 review from a tidy workbook, with and without the ModelCensus toolkit prompt. Every figure on every deck is checked against the workbook in code.
The same request and the same toolkit prompt on workbooks shaped like real finance exports: subtotals that include intercompany, negatives in parentheses, a renamed region, units in a note, blank cells for a new region.
On messy enterprise data, what buys more accuracy per dollar: describing the data, or buying a bigger model?
A refund agent meets an injected instruction. Does a permission gate stop it, or labels on what the agent may trust? Which plane stops which failure?