What we measure, what we refuse to measure, and why no per-model reliability finding is published today.
Current status: no per-model reliability finding is published. The publishing path is closed in source control, not merely unused.
We assess how complete our own governance record is for each benchmark. We do not score benchmarks, and we do not publish per-model reliability findings. The machinery for the second thing is built and tested; the switch that would turn it on is a constant in source control, and it is off.
For each benchmark we record five governance facts: what contamination controls the maintainers describe, how often the set is updated, its licence, how it is accessed, and whether it is saturated. Coverage says how many of those five we hold.
This is a fact about us, not a grade on the benchmark. Anyone can check it against the register. A complete record does not mean the benchmark is reliable, and we never say that it does; an incomplete record does not mean the benchmark is weak. The reader draws whatever inference they think the record supports. We publish the record.
There are three bands because a record has exactly three states: empty, partial, complete. There is no number attached to any of them. A coverage percentage would be a certainty dial with the label removed, and research on decision support found that pairing hedged wording with a numeric confidence measurably increases reliance on wrong recommendations.
Where a benchmark has documented incidents, we count them by kind and show them next to the coverage band. We do not subtract them from it. A benchmark with a detailed incident log is usually one whose maintainers disclose problems, and an index that quietly penalised disclosure would reward silence instead.
None of these is published today. They are recorded here so that the bar is visible before anything meets it, and so that a reader can hold us to it later.
Membership inference and perplexity based detectors are excluded, and not as a matter of caution. A 2026 replication across several large models found these methods performing at roughly the level of chance, and separate work found that baselines ignoring the model entirely matched or beat the published state of the art. A signal at chance level cannot support a public statement about a named product, so it is absent from our tier vocabulary rather than merely discouraged.
We also do not combine several weak detectors into a single index. Averaging uncertain signals produces a confident looking number without producing any more evidence, and it is the same certainty dial in a different costume.
We use licensed access or permissively licensed public data. We do not rotate credentials or disguise requests to work around a rate limit or a ban, under any circumstances.
Each version of this document is fixed once published. We do not edit a published version in place, because an assessment stored against a version is only meaningful if that version still means what it meant. A change to the method or to this text is a new version with a changelog entry, and earlier assessments keep their original version label.
If we get something wrong we correct it in the open: the correction is added, the original stays readable, and nothing is quietly removed.
Initial version. Governance coverage bands, the evidence bar for a per-model finding, and the exclusions. No finding is published.