ModelCensusopen-source ai reliability harness

Reliability engineering, for language models.

Every engineering discipline has a failure taxonomy. Generative AI doesn’t.

Not a leaderboard. An open census of how models behave, where they fail, and what they’re actually fit to do. Every number opens to the call behind it.

Replay an experiment →

Open taxonomy & harnessDeterministic detectorsReplayable labs

40 published loops · 20 models · 10 failure modes · 5 labs

Class 1 · Grounding & Attribution

Making It Up

Problem: Confident claims and citations with nothing real behind them.

Engineering guidance: Supply the source rather than asking the model to recall one — a real passage substantially reduces post-cutoff fabrication across the panel.

See an example →Modes in this class →

Grounding &AttributionSycophancy &EpistemicIntegrityInstructionAdherence &Long-ContextReasoning &CalculationRobustness &ConsistencyTool Use &Agentic ControlTemporal &KnowledgeBoundaryHow LLMsfail7 classes
Context sensitivity

Context is not a general-purpose fix.

Supplying a source is everyone’s fix for unreliable answers. It works on some failures, does nothing for others, and makes a few worse.

Each axis is a failure mode: the dotted line is the model alone, the shape is the same model with a source. Know which failures context can reach before you build retrieval for them.

Every model, every mode →What each axis means →

Context sensitivity by failure mode
0–100% failureNo contextbaselineCites sourcesthat don't existc1 · 35% [30–40]Invents where anumber came fromc2 · 24% [19–30]States facts itcan't knowc7 · 17% [13–22]Forgets a ruleyou setc3 · 16% [10–25]Breaks the JSONshapec3 · 8% [4–15]Right method,wrong mathsc4 · 0% [0–1]Right number,wrong unitc4 · 0% [0–1]Answer changeswith wordingc5 · 1% [0–3]Loses the runningtotalc6 · 4% [2–6]Gets dates wrongc7 · 11% [8–16]

band = middle 80% of 20 models · 10 modes · set s1.4-live · spans 7 of 7 classes · axes are modes, never classes

Dotted = no-context baseline, fixed. The gap is the impact. Further out = fails more.

Cost per correct answer

The meter rate isn't the fare.

Price per token is the meter. A right answer is the fare, and the data around the model sets it as much as the model does.

In one lab, the cheapest model with a short data dictionary beat a frontier model on raw tables: more answers right, at 44× less per correct answer.

Replay the experiment →Tokenomics is an architecture decision →

GPT-6 Lunacheapest modeldictionary~150 tokens91% right1× per right answerClaude Opus 5.5frontier modelmessy datapr 1 · dq Y80% right44× per right answer
Weight is right answers. From the semantic-layer lab: 15 questions over messy tables, four models, 540 recorded calls.
Information gain

More context isn't more signal.

More tokens don’t mean better answers. Irrelevant context costs money and accuracy.

So the census holds the amount of context fixed and varies only its quality. A change can then be pinned on what the context says, not how much of it there is.

How the four arms work →The control most evals skip →

Low qualitynoisy · unfilteredHigh qualitycurated · denseHigh quantitylong windowLow quantityshort windowThe Distraction ZoneHigh latency, context rot,negative information gain.The Power EngineMax throughput, high signal,true synthesis.The Blind SpotIncomplete answers, parametrichallucination.The Sweet SpotPrecision answers, lowestcost, peak efficiency.This census measures the bottom row.volume held constant · every arm 160–175 tokens
The panel

Model performance is a moving target.

One failure mode ranges from 0% to 100% across today’s models. Newer isn’t safer, and a model ID doesn’t pin the model: routing, quantization and silent updates move it.

So every run records the model that actually answered, and a re-run is a new measurement, not a confirmation.

Model by model →Why an id is not a model →

Constraint Decay at Depth20 models · same cases · same harness0%25%50%75%100%newest · 2026oldest · 2023Same question set. 0% to 100% failure.each dot is one model · newest at the top · rates in the studio
Labs

Experiments you can replay.

Recorded once against the live APIs, every call kept. Pick a case, press replay, and watch the models answer.