VERIK / V106 / 28 JUL 2026
Operating in the FogAcademic

The Ceiling on What a Red-Team Result Can Prove

On July 23, 2026, a 21-page paper appeared on the arXiv preprint server that asks a question the current wave of frontier-model governance frameworks has not asked out loud. The paper is What AI Red-Team Evaluations Can and Cannot Prove by Bandana Kaur. Its central move is to define an object the paper calls the evidential ceiling of an evaluation, derive it in closed form, and then audit eight publicly available evaluation suites against the boundary the ceiling implies.

The finding is narrow, technical, and unusually direct. Above a calculable harm rate, a benchmark of modest size can certify a category to a stated evidentiary standard, and a clean sheet is the stronger of the two possible observations. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the assumed scoring rule and trial structure. The crossing between the two regimes has a closed form. The bound is not tied to any one benchmark. Written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the two hypotheses, not attack success, is what determines evidential worth.

The eight suites the paper audits are, on the author's telling, adequate for high-frequency harm categories and several orders of magnitude short for the rare catastrophic categories that the frontier-model governance discussion most often invokes.

That is the paper. Its structural implication is what makes it a load-bearing artifact for this week in particular.

Three institutions announced evaluation instruments on the same week

The paper landed in a week in which three separate institutional actors named evaluation as the operational core of their frontier-model governance posture.

On July 27, the National Institute of Standards and Technology announced the AI Technology Evaluation platform, an isolated testbed for image analysis in quantum science, genomics, and public safety, with first evaluations scheduled for August. On the same day, the National Institute of Standards and Technology published the initial public draft of Special Publication 800-239, which frames AI data center facilities themselves as governance objects with a comment period through September 25. On the same day, reporting on the near-final White House voluntary frontier AI review framework confirmed that the framework's core mechanism is a classified benchmarking process operated by the National Security Agency to designate which models are covered frontier models, coupled to Center for AI Standards and Innovation pre-deployment evaluations already signed by five of the six major frontier labs.

Each of these instruments is an evaluation instrument. Each of them asks the same underlying question the Kaur paper takes as its subject. Each of them, under the paper's ceiling result, sits at the top of a curve whose shape has now been drawn.

What the ceiling actually says

The paper's contribution is not a claim that evaluations are useless. Its final line states the opposite. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions.

The contribution is a boundary condition. For a fixed testing budget and a fixed scoring rule, the largest factor by which a single benchmark result can move a prior belief is calculable. When the underlying harm rate the evaluator is trying to bound is high enough, the ceiling is above the evidentiary standard the governance regime requires, and a clean pass is a meaningful outcome. When the harm rate is low enough, the ceiling is below the evidentiary standard, and no passive benchmark of feasible size produces the evidence the standard asks for.

The paper generalizes from benchmarks to adaptive procedures. The elicitation rate under the null hypothesis, not the elicitation rate under the alternative, is what determines the ceiling. Automated red teams that maximize attack success without also maximizing discrimination between hypotheses do not improve the ceiling. They improve the number of attacks found, which is a different quantity.

What the frontier framework is actually asking evaluations to do

The frontier-model governance frameworks now under construction ask evaluations to do two things. They ask evaluations to characterize known capability, which is the high-frequency regime the paper says benchmarks handle. They also ask evaluations to bound the probability of catastrophic misuse, which is the low-frequency regime the paper says benchmarks of feasible size do not reach.

The classified benchmarking process the National Security Agency has been directed to finalize by the August 1 statutory deadline is being architected to produce a designation. The designation, under the reported terms of Executive Order 14409, triggers a 30-day voluntary early-access window during which agencies including the National Security Agency and the Treasury Department may examine the model before broader release. The classification of a model as covered is meant to carry information about the model's frontier-capability profile relative to national security concerns.

Under the ceiling result, the confidence that a classified benchmark can produce in the low-frequency regime is bounded by the same closed-form quantity that bounds any other benchmark of feasible size. Classifying the benchmarking process does not change the evidentiary standard the process must satisfy. It changes who can inspect the result.

The gap the paper names

The gap the paper names is not a gap between careful evaluation and careless evaluation. It is a gap between what the evaluation instrument can prove and what the surrounding governance instrument reports the evaluation to have proved. When a covered-frontier designation, a Center for AI Standards and Innovation attestation, or an AI Technology Evaluation report enters a downstream governance record, the record does not carry with it the ceiling of the underlying benchmark. The ceiling stays with the paper on the arXiv preprint server. The record travels alone.

The discipline the paper says the field needs is a discipline of stating which propositions a benchmark is informative about. That discipline is not a property of the models under evaluation, or of the evaluators, or of the classification level of the benchmarking process. It is a property of the way the evaluation result is written into the governance instrument.

What remains on the table:

The loop closed around an oversight function that was never instrumented.