VERIK / V110 / 03 AUG 2026
Operating in the FogGovernance

The Evaluation Floor Now Moves With Compute

On July 28, 2026, OpenAI researchers submitted GPT-Red: Automated Red Teaming via Self-Play at Scale to arXiv. The paper describes an automated adversary trained by reinforcement learning against a diverse population of simultaneously-trained defender models, using compute on the same order as the vendor's largest reinforcement learning post-training runs. The authors describe the training run, in the abstract itself, as "the single-largest LLM safety training run ever documented."

That characterization is the load-bearing sentence. It is not a claim about a particular attack, or a particular benchmark, or a particular model. It is a claim about where the red-team floor now sits.

In the replicated indirect prompt injection arena of Dziemian et al. 2025, GPT-Red achieved an 84 percent attack success rate on held-out scenarios. Human red-teamers in the same replication reached 13 percent. Against GPT-5.5, GPT-Red reliably breaks the model. Against GPT-5.6 Sol, the vendor's adversarially trained defender, GPT-Red's own direct injections fail 0.05 percent of the time. A single attack class the paper introduces, described as "fake chain of thought," fell from above 95 percent success on GPT-5.1 to below 10 percent on GPT-5.6 Sol in the four-month interval covered by the paper. These are first-party numbers. The attacker is not released. The defender is deployed. Independent replication of any of it is not possible with the artifacts made public.

What the Instrument Now Measures

An adversary trained at post-training compute scale is a different evaluation object from a human red team. It is priced differently. It moves on a different clock. It generalizes across defender models and harnesses because it was trained to. It does not tire. It does not require scheduling. It becomes stronger as the defender does, because the paper describes a self-play flywheel in which each generation of defender provides a better learning signal for the next generation of attacker.

The academic and regulatory instrument that governs pre-deployment safety in most jurisdictions today is the red-team evaluation, understood as an artifact produced by qualified adversarial testers within a bounded window before release. That instrument was designed against a human-red-team baseline. It answers the question: given a team of experts and a defined window, what did they find? The report is what evaluators found under the conditions they tested, with the elicitation techniques they possessed, inside the containment architecture they built. It is a documented lower bound.

GPT-Red does not falsify that instrument. It changes what the lower bound is a lower bound of.

If an internal automated adversary can be trained at the scale of a frontier post-training run and can generalize to held-out defenders, then a "passing" evaluation now records, in effect, what the last-generation adversary could not elicit at the time. Absent from the record is what the current-generation adversary could elicit, or what next-generation adversary will elicit six months from now on the same model at the same deployment level. The evaluation window closes. The compute frontier does not.

The Kaur Ceiling as a Structural Constant

A week earlier, on July 22, 2026, arxiv.org/abs/2607.21735 by Bandana Kaur formalized the epistemic scope of red-team evaluations: what they can and cannot prove. Kaur's argument is that a red-team result documents a lower bound on capability under specified conditions, and only under those conditions. A passing evaluation is evidence that dangerous capability was not elicited by the techniques applied. It is not evidence that dangerous capability is absent.

GPT-Red is the empirical instantiation of that structural claim. The bound Kaur described in the abstract now has an artifact attached to it. The techniques applied at time T are no longer the techniques an evaluator with a comparable compute budget could apply at time T plus one training cycle. What was elicited was what the elicitation apparatus of a specific moment could reach. The gap between what the apparatus reached and what a later apparatus reaches is the gap between the retained evaluation report and the current safety posture of the deployed model.

The HANDBOOK.md benchmark submitted July 28, 2026 by Panavas and coauthors described the same shape at the standing-instruction layer: a policy artifact retained in context that decays as the operative control on agent behavior. Three governance-instrument frames in one window describe three layers of the same problem. The instrument is retained. The function bounded not by the license or the guideline or the report, but by structural features of the object the instrument is written against.

What the Vendor Retains

OpenAI does not publish executable attack prompts. The offensive model is not distributed. The vendor's stated reason is proliferation risk. The stated compute figure, described in the paper as comparable to the vendor's largest post-training runs and elsewhere characterized as roughly 700,000 GPU hours, is a self-report. The stated results are self-reports. The Dziemian et al. arena replication is described in the paper. The vending-machine case study and the Codex CLI exfiltration case studies are described in the paper. Both live agents were attacked; the vending-machine attack included fabricated administrator metadata and reportedly changed prices, placed unauthorized orders, and cancelled a customer's order in the disclosed environment.

An external party wishing to verify any of it has access to the abstract, the paper, the model behaviors, and the assurances of the vendor. The apparatus that produced the numbers is retained by the vendor. The apparatus that a third party would need to reproduce those numbers is not distributed.

The evaluation record is therefore doubly bounded. It is bounded by what the internal adversary at the moment of publication could elicit, and it is bounded by the visibility of the elicitation apparatus itself. A regulator reading the abstract has the vendor's statement. A regulator wishing to check the statement has the same statement.

What Remains on the Table

The loop closed around an oversight function that was never instrumented.