ModelRiskIndex

Methodology v0.1 · 2026-08-04

Nothing is a black box

ModelRiskIndex grades production AI models on five things a buyer must answer, rolls them into a tier, and links every grade to evidence. This page is the standard; it has a version number and a changelog because methodology changes are changes too. Disagreements are expected — vendors get a formal rebuttal channel, and rebuttals are published verbatim alongside the contested grade.

The five vectors

Five questions, in the order they decide a deal. Each is graded Strong, Partial, or Weak. Two further states are first-class and distinct from a weak grade: not applicable marks vectors that are properties of the deployer rather than the graded artifact (common for open weights), and under review marks vectors we have not yet assessed or where evidence is in dispute. A grade is a human judgment citing underlying numbers — third-party scores are never averaged across sources, because their methodologies differ.

The framework is deliberately capped at five. Adversarial resistance folds two measures — direct jailbreaks and agentic prompt injection — into one graded slot, shown to the weaker of the two because an attacker takes the easier path; both remain visible as facets on every model page. The test for any proposed sixth vector: does it answer one of these five questions better? If so it is detail that belongs inside one of them. If it is a genuinely new question, only then does the wheel grow.

Tiers

A model's tier is computed, never assigned: each model answers a fixed requirements checklist, and the tier — along with the exact list of what is missing for the next tier — is derived from those answers. Every checklist answer is visible on the model page. Adding a requirement to this standard forces every tracked model to answer it before the dataset will build; the bar can be raised, but never silently.

A requirement can be marked not-applicable (with justification) when it cannot meaningfully apply to the artifact — e.g. enterprise data controls for self-hosted open weights. Not-applicable counts as satisfied but is always rendered as N/A, never as a pass.

Composite score

The 0–100 composite is deliberately simple: Strong = 2, Partial = 1, Weak = 0, summed across the applied grades and normalized. Vectors marked not-applicable or under-review are excluded from both numerator and denominator. It exists as a sort key, not a truth claim — the vectors and their evidence are the product.

How our own claims are verified

A site that grades others on transparency has to be checkable itself. Three mechanisms, in increasing order of strength:

Quote binding is being backfilled across the registry; the verification endpoint publishes the current coverage honestly rather than waiting for it to be complete. An unbound claim is unverified, which is not the same as wrong — the same distinction we apply to the models we grade.

What this cannot do: verify third-party measurements (we can confirm we reported a number correctly and cite its methodology, not that the methodology is sound), or verify a judgment. Grades are human judgments citing evidence; the evidence is verifiable, the judgment is arguable, and keeping those separate is deliberate.

Evidence rules

Open-weight models

Open-weight models are graded on a self-hosted reference deployment: data handling and compliance are deployer properties, so they are marked not-applicable rather than graded, while jailbreak and injection resistance are graded on the weights as released, without optional guardrail layers. Hosted first-party endpoints of open-weight models (e.g. DeepSeek's API) are graded as what they are: hosted services with their own terms. The Mistral objection — that deployers control fine-tuning and guardrails — has merit, and this split is our current answer to it.

Conflicts and corrections

When sources disagree materially, the model page shows both with dates and links, and the grade follows documented precedence with the disagreement noted. Corrections ship as visible dataset changes with the error acknowledged — no silent fixes, since the entire premise of this site is that silent changes are bad.

What we are not

Rankings are free and will not be paywalled. No lab pays for placement. Any future commissioned assessment work will be structurally separated from public scores. Revenue comes from monitoring alerts, procurement exports, and data licensing — from buyers, not from the labs being graded.

Methodology changelog