Measure how models fail — and prevent failures you can't see.
ModelCensus combines deterministic failure-mode evaluations with real-time interceptors to stress-test your models, expose blind spots, and block unseen errors before they reach your users. Every rate ships with a 95% confidence interval and the transcript behind it — and the CFMI reports which models it cannot tell apart.
- Modes instrumented
- of 27 catalogued
- Models in panel
- fixed, pre-registered
- Scored trials
- excludes unverifiable
- Failure classes
- grounding → temporal
no ranking — you compare cells
The census sheet
One control chart per failure mode, seven rows by class. Nineteen independent processes with independent limits — a layout, not an aggregation.
One chart per failure mode, bare condition. Each is its own process with its own limits — nothing is pooled across modes, and there is no total. 1 wave so far; control limits appear at 8.
| Mode | Wave | Failure rate | n | Signal |
|---|---|---|---|---|
| fmi_1_2 Citation Resolution Failure | showcase-2026-q3 | 62.1% | 29 | none |
| fmi_1_3 Claimed-Action / Tool-Log Divergence | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_1_4 Planted-History Recall Failure | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_2_1 Pushback Instability | showcase-2026-q3 | 46.7% | 15 | none |
| fmi_2_3 Authority-Framing Sensitivity | showcase-2026-q3 | 0.0% | 14 | none |
| fmi_2_4 Basis-Demand Evasion | showcase-2026-q3 | 33.3% | 15 | none |
| fmi_3_1 Constraint Decay at Depth | showcase-2026-q3 | 20.0% | 15 | none |
| fmi_3_3 Cross-Turn Schema Validity | showcase-2026-q3 | 6.7% | 15 | none |
| fmi_3_4 Needle-Position Sensitivity | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_4_1 Method/Execution Split | showcase-2026-q3 | 3.3% | 30 | none |
| fmi_4_4 Dimensional-Analysis Failure | showcase-2026-q3 | 20.0% | 15 | none |
| fmi_5_2 Paraphrase Non-Invariance | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_6_1 Tool-Routing Error | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_6_2 Tool-Argument Validity | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_6_3 Agentic Loop / Non-Termination | showcase-2026-q3 | 6.7% | 15 | none |
| fmi_6_4 Tool-Return Groundedness | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_6_5 Ledger Reconciliation Failure | showcase-2026-q3 | 0.0% | 15 | none |
| fmi_7_1 Post-Cutoff Fabrication | showcase-2026-q3 | 33.3% | 15 | none |
| fmi_7_3 Date-Arithmetic Error | showcase-2026-q3 | 33.3% | 15 | none |
CFMI — cumulative failure index
Severity-weighted mean failure rate, lower is better; bars show the 95% interval. Models are listed alphabetically — this is not a ranking.
No census is published yet — the index appears here once a run passes review. The taxonomy and the method are readable now.
The seven failure classes
Every mode belongs to exactly one class. Classes group failures by mechanism, not by severity.
Confident claims and citations with nothing real behind them.
Abandoning a correct answer when a user pushes, flatters, or name-drops.
Constraints set early that quietly evaporate under length and distraction.
Right method, wrong execution — and the commonsense gotchas that pattern-matching walks straight into.
Same question, reworded — and the answer changes.
Wrong tool, invalid call, endless loop, or an answer ungrounded in what the tool returned.
Stale facts stated as current, and date math against a world it can't see.
How to read this
Two ways in, depending on what you own.
A fixed, versioned case is sent to a model as served that day, with the serving provider recorded. A deterministic detector scores the response as PASS, FAIL, NOT_APPLICABLE or ERROR. Only PASS and FAIL enter the denominator: the rate is FAIL divided by PASS plus FAIL, reported with a 95% Wilson confidence interval. Counting an unverifiable check as a failure would measure publisher policy rather than the model.
Evaluator libraries give you machinery. None tells you which failures to look for, or how often frontier models commit them. That is the layer here: a named taxonomy, a detector per mode, published population rates to compare your own results against.
Where FMI sits alongside LangSmith, Inspect and HELM →Generation got cheap; verification did not. Every output entering a decision path creates a checking obligation, absorbed by whoever is senior enough to catch a confident error. A failure rate is the rate at which that happens — and the modes hardest to notice cost the most, which is why stealth is weighted alongside harm.
How the rates are measured →