ModelCensusfailure-mode benchmark
The AI Failure Mode Census

Measure how models fail — and prevent failures you can't see.

ModelCensus combines deterministic failure-mode evaluations with real-time interceptors to stress-test your models, expose blind spots, and block unseen errors before they reach your users. Every rate ships with a 95% confidence interval and the transcript behind it — and the CFMI reports which models it cannot tell apart.

Deterministic detectorsEvery published exhibit human-reviewedOpen taxonomy · open harness
Current censusunpublished
Modes instrumented
of 27 catalogued
19
Models in panel
fixed, pre-registered
Scored trials
excludes unverifiable
Failure classes
grounding → temporal
7

no ranking — you compare cells

The census sheet

One control chart per failure mode, seven rows by class. Nineteen independent processes with independent limits — a layout, not an aggregation.

How the limits work

One chart per failure mode, bare condition. Each is its own process with its own limits — nothing is pooled across modes, and there is no total. 1 wave so far; control limits appear at 8.

Failure rate by mode and wave, bare condition
ModeWaveFailure ratenSignal
fmi_1_2 Citation Resolution Failureshowcase-2026-q362.1%29none
fmi_1_3 Claimed-Action / Tool-Log Divergenceshowcase-2026-q30.0%15none
fmi_1_4 Planted-History Recall Failureshowcase-2026-q30.0%15none
fmi_2_1 Pushback Instabilityshowcase-2026-q346.7%15none
fmi_2_3 Authority-Framing Sensitivityshowcase-2026-q30.0%14none
fmi_2_4 Basis-Demand Evasionshowcase-2026-q333.3%15none
fmi_3_1 Constraint Decay at Depthshowcase-2026-q320.0%15none
fmi_3_3 Cross-Turn Schema Validityshowcase-2026-q36.7%15none
fmi_3_4 Needle-Position Sensitivityshowcase-2026-q30.0%15none
fmi_4_1 Method/Execution Splitshowcase-2026-q33.3%30none
fmi_4_4 Dimensional-Analysis Failureshowcase-2026-q320.0%15none
fmi_5_2 Paraphrase Non-Invarianceshowcase-2026-q30.0%15none
fmi_6_1 Tool-Routing Errorshowcase-2026-q30.0%15none
fmi_6_2 Tool-Argument Validityshowcase-2026-q30.0%15none
fmi_6_3 Agentic Loop / Non-Terminationshowcase-2026-q36.7%15none
fmi_6_4 Tool-Return Groundednessshowcase-2026-q30.0%15none
fmi_6_5 Ledger Reconciliation Failureshowcase-2026-q30.0%15none
fmi_7_1 Post-Cutoff Fabricationshowcase-2026-q333.3%15none
fmi_7_3 Date-Arithmetic Errorshowcase-2026-q333.3%15none

CFMI — cumulative failure index

Severity-weighted mean failure rate, lower is better; bars show the 95% interval. Models are listed alphabetically — this is not a ranking.

Full matrix

No census is published yet — the index appears here once a run passes review. The taxonomy and the method are readable now.

The seven failure classes

Every mode belongs to exactly one class. Classes group failures by mechanism, not by severity.

Full index

How to read this

Two ways in, depending on what you own.

The framework
How one number is made

A fixed, versioned case is sent to a model as served that day, with the serving provider recorded. A deterministic detector scores the response as PASS, FAIL, NOT_APPLICABLE or ERROR. Only PASS and FAIL enter the denominator: the rate is FAIL divided by PASS plus FAIL, reported with a 95% Wilson confidence interval. Counting an unverifiable check as a failure would measure publisher policy rather than the model.

If you build evals

Evaluator libraries give you machinery. None tells you which failures to look for, or how often frontier models commit them. That is the layer here: a named taxonomy, a detector per mode, published population rates to compare your own results against.

Where FMI sits alongside LangSmith, Inspect and HELM →
If you own the estate

Generation got cheap; verification did not. Every output entering a decision path creates a checking obligation, absorbed by whoever is senior enough to catch a confident error. A failure rate is the rate at which that happens — and the modes hardest to notice cost the most, which is why stealth is weighted alongside harm.

How the rates are measured →