Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

What AI Red-Team Evaluations Can and Cannot Prove

arXiv:2607.217352026-07-22

There's a calculable limit to how much a 'no harm found' AI safety test can actually prove

This paper works out, in closed-form math, exactly how much evidence a red-team safety test can provide when it finds zero harmful outputs. For common harm categories (around a 1 percent rate), existing public benchmarks already provide meaningful certification, but for rare categories, no benchmark of any feasible size can certify safety at all past a calculable boundary. The authors argue labs should compute this boundary before running an evaluation, not after.

METAL MEDIA explanatory visual

How the evidential ceiling decides what a safety test can and can't prove

Evidence statusMeasured results reported

  1. Set up the testDefine the safe vs. unsafe hypothesis pair, sample size n, target improvement ratio r, and acceptable risk threshold tau
  2. Score the clean resultCompute the likelihood ratio for zero harmful outputs to see exactly how much it shifts belief toward safety
  3. Locate the boundary pminTheorem 1: above pmin, a finite feasible sample size can certify safety. Theorem 2: below pmin, no feasible benchmark size can
  4. Check real benchmarks against itSimulate eight suites (AdvBench, HarmBench, SafetyBench, etc.) at 1 percent and 0.0001 percent harm rates to measure their statistical power and false-certification risk
  5. Standardize reportingRequire sample size, clustering info, and confidence intervals to be published together so readers know exactly what claim a result supports
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. The paper defines an 'evidential ceiling': the maximum factor by which a single test result can shift belief about whether a model is safe, given a fixed testing budget.
  2. It derives a closed-form formula for how much evidence a clean benchmark result (zero harmful outputs observed) actually carries, letting anyone compute the minimum sample size needed to certify safety at a given harm rate and improvement target.
  3. At a 1 percent harm rate, existing public benchmarks like AdvBench (520 prompts) can already support a tenfold shift in belief toward safety, but at a 0.1 percent rate the required sample jumps to 9,203 prompts, and at 0.01 percent to 92,096.
  4. Below a specific boundary harm rate (about 1.4x10^-5 under one set of assumptions: Nmax=10^5, r=0.5, tau=0.5), the paper mathematically proves that no benchmark of any feasible size can turn a clean result into meaningful evidence of safety.
  5. Monte Carlo simulation (8,000 iterations per point) of eight evaluation suites (AdvBench, HarmBench, SafetyBench, XSTest, etc.) found that at a 1 percent harm rate, none reached 80 percent statistical power, while most suites were adequately powered at an 8 percent harm rate.
Figure 1: The evidential ceiling and the two evidence regimes. (a) The evidence channel. Both ceilings are fixed by the channel and the budget; by the data processing inequality no rescoring or aggregation can raise either. The exculpatory ceiling Ceil− governs certification, and for a passive benchmark it is attained at k=0. (b) Evidence contributed by a single result, in bits, both quantities scored against the same hypothesis pair (H0: p=r​pu versus H1: p=pu) at r=0.5. Solid curves are a clean sheet, dashed curves one observed harmful output. Open circles mark the crossing rate of Corollary 1. To the right of a circle the clean sheet is the stronger evidence; to the left it is not. As p falls the clean sheet carries vanishing evidence while the single harm converges to log2⁡(1/r)=1 bit, independent of n.
Figure 1: The evidential ceiling and the two evidence regimes. (a) The evidence channel. Both ceilings are fixed by the channel and the budget; by the data processing inequality no rescoring or aggregation can raise either. The exculpatory ceiling Ceil− governs certification, and for a passive benchmark it is attained at k=0. (b) Evidence contributed by a single result, in bits, both quantities scored against the same hypothesis pair (H0: p=r​pu versus H1: p=pu) at r=0.5. Solid curves are a clean sheet, dashed curves one observed harmful output. Open circles mark the crossing rate of Corollary 1. To the right of a circle the clean sheet is the stronger evidence; to the left it is not. As p falls the clean sheet carries vanishing evidence while the single harm converges to log2⁡(1/r)=1 bit, independent of n.
Table 1: Two regimes, one hypothesis pair. Evidence in bits carried by a clean benchmark and by one observed harmful output, both scored against H0: p=r​pu versus H1: p=pu, at n=520 and r=0.5. The ordering reverses at p×=1.33×10−3 (Corollary 1). Neither observation is universally the stronger.
pu10−25×10−310−310−410−5
|log2⁡Λ0|, clean sheet3.7791.8820.3750.0380.004
|log2⁡Λ1|, one observed harm2.7720.8820.6250.9630.996
ratio, harm to clean sheet0.730.471.725.7265.6
Figure 2: Statistical power across eight evaluation suites. Power to detect a fifty percent reduction in harm rate, from 8,000 Monte Carlo iterations per operating point at α=0.05. Dashed lines mark the frontier operating point (p=0.01) and a high-frequency category (p=0.08). At p=0.08 most suites are adequately powered; at p=0.01 none reaches 80 percent.
Figure 2: Statistical power across eight evaluation suites. Power to detect a fifty percent reduction in harm rate, from 8,000 Monte Carlo iterations per operating point at α=0.05. Dashed lines mark the frontier operating point (p=0.01) and a high-frequency category (p=0.08). At p=0.08 most suites are adequately powered; at p=0.01 none reaches 80 percent.
Table 2: Discrimination, not attack success, determines evidential worth. Evidence in bits carried by a null result, for illustrative hypothesis-conditioned rates. The values are chosen to display the structure and are not empirical estimates for any published procedure. Rows three and five have the same elicitation rate under H1 and differ by two orders of magnitude in the evidence a null result carries.
ProcedureTrial unitnq1q0bits
Passive benchmark, AdvBench-sizedprompt52010−35×10−40.38
Passive benchmark, at Nmaxprompt10510−35×10−472.19
Adaptive campaign, discriminatingcampaign10.900.103.17
Adaptive campaign, discriminatingcampaign50.900.1015.85
Adaptive campaign, non-discriminatingcampaign50.900.852.92
Figure 3: False certification and the boundary. (a) False certification against benchmark size for four harm rates, with the range of current public benchmarks shaded. (b) The boundary pmin​(τ,Nmax,r) against the feasibility ceiling, for three combinations of τ and r. Above a curve, Theorem 1 gives a finite sufficient n; in the shaded region below, Theorem 2 applies and no feasible benchmark certifies.
Figure 3: False certification and the boundary. (a) False certification against benchmark size for four harm rates, with the range of current public benchmarks shaded. (b) The boundary pmin​(τ,Nmax,r) against the feasibility ceiling, for three combinations of τ and r. Above a curve, Theorem 1 gives a finite sufficient n; in the shaded region below, Theorem 2 applies and no feasible benchmark certifies.
Table 3: False-certification probability across the operating surface. The probability that a model with true harm rate p produces zero harmful outputs in n approximately independent trials. Values above roughly five percent mark operating points at which a clean result is not informative evidence. Current public benchmarks occupy the leftmost three columns, where they are adequate at p=10−2 and inadequate below it.
P⁡(k=0∣n,p), %n=250n=520n=2,100n=5,000n=104n=105
p=10−28.10.50.00.00.00.0
p=10−377.959.412.20.70.00.0
p=10−497.594.981.160.736.80.0
p=10−599.899.597.995.190.536.8
Figure 4: Benchmark and deployment prompt distributions. Joint UMAP projection [25] of AdvBench, HarmBench, and a 9,089-query sample of LMSYS-Chat-1M under sentence-transformer embeddings. The benchmarks occupy a narrow region measurably separated from ordinary deployment traffic. This does not bound the distance to the adversarial component of deployment, which is the component catastrophic-risk claims concern.
Figure 4: Benchmark and deployment prompt distributions. Joint UMAP projection [25] of AdvBench, HarmBench, and a 9,089-query sample of LMSYS-Chat-1M under sentence-transformer embeddings. The benchmarks occupy a narrow region measurably separated from ordinary deployment traffic. This does not bound the distance to the adversarial component of deployment, which is the component catastrophic-risk claims concern.
Table 4: The claims ladder. Required sample size, at r=0.5 and one-sided α=0.05, for each level of claim at four harm rates. Current public benchmarks (n≤2,100) support every claim in the table at p=10−2, the weakest two at p=10−3, and only the upper bound below that. The final row applies Theorem 2 at Nmax=105.
Claim the evaluation licensesp=10−2p=10−3p=10−4p=10−5
Harm rate bounded above at 95% confidenceany nany nany nany n
Belief shifts toward safety by 2× (τ=0.5)1381,38513,862138,628
False certification held below 5%2982,99429,956299,572
Belief shifts by 10× (τ=0.1)4574,60246,048460,514
Belief shifts by 100× (τ=0.01)9149,20392,096921,027
Feasible at Nmax=105?yesyespartlyno
Table 5: Minimum reporting template for a red-team null result. To be completed once per harm category, per model, before the result enters a safety case.
FieldSymbolRequirement
Harm categoryNamed, with the threat model it operationalizes
Elicitation procedurePassive corpus, adaptive campaign, or other; stated
Trial unitPrompt, campaign, or replication; campaigns are not prompts
Nominal sample sizenExact count of independent trials in the stated unit
Event countkExact integer, at the same analysis unit as n
Generations per promptmExact; state if greater than one
Intra-cluster correlationρEstimated from the run, or an upper bound
Effective sample sizeneffn/[1+(m−1)​ρ]
Upper confidence boundExact Clopper-Pearson at 95% on the nominal (n,k)
Clustering-adjusted boundBeta-binomial or GEE; not Clopper-Pearson on neff
Minimum detectable effectδminPrespecified, not chosen post hoc
Observed power1−βAgainst δmin, with sidedness stated
Evidentiary thresholdτWith the loss ratio L and prior odds O0 that fix it
Elicitation rate under H1q1Measured against a positive control; state the control
Elicitation rate under H0q0Stated or bounded; q1 alone is not sufficient
Per-trial discriminationκ|log2⁡[(1−q1)/(1−q0)]|, in bits
Achieved likelihood ratioΛ0[(1−q1)/(1−q0)]n exactly, reported as a number
Claim licensedThe highest row of Table 4 the design supports
Distributional coverageDistance to each deployment mixture component
Semantic consistencyσAcross prompt transformation types
Out-of-distribution gapAttack success on held-out attack families
Position relative to boundaryWhether p is suspected below pmin​(τ,Nmax,r)

Findings

  • At a 1 percent harm rate with 520 prompts (AdvBench-sized), the paper calculates that a clean result carries 1.4 times more evidence than a single observed harmful output, with the crossing point between the two occurring at a harm rate of about 1.33x10^-3.
  • Power simulations across eight evaluation suites (AdvBench, HarmBench, SafetyBench, XSTest, StrongREJECT, etc.) showed power ranging from 14.9 percent (XSTest) to 59.3 percent (SafetyBench) at a 1 percent harm rate, all below the 80 percent adequacy threshold.
  • At a 0.1 percent harm rate, StrongREJECT (313 prompts) has a 73.1 percent chance of falsely certifying an unsafe model as safe, and AdvBench has a 59.4 percent chance.
  • AdvBench and HarmBench prompts sit 3.1 to 3.5 times farther from real user conversation data (a 9,089-query sample from LMSYS-Chat-1M) in embedding space than the benchmarks are from each other, showing they don't represent ordinary deployment traffic.
  • Reviewing nine public frontier model safety disclosures, the authors found only one endpoint reported both an exact numerator and denominator for a binary harm rate, and none reported the dependence structure needed to interpret it.

Where it can be used

  • AI labs could use this framework to check, before running an evaluation, whether a benchmark of a given size can even in principle certify safety for a specific harm category at a stated confidence level.
  • Auditors or regulators could use the same math as a checklist to interpret what a 'zero harmful outputs observed' line in a system card actually implies statistically.
  • Teams writing safety reports could adopt the paper's minimum reporting template (sample size, clustering structure, confidence interval) to make their claims verifiable.

Limits and open work

  • The core mathematical inequality behind both theorems dates back to 1983; the paper's contribution is applying it to evaluation design, not the math itself.
  • The harm-rate estimates used in the power analysis are described as defensible rather than authoritative, since published estimates for frontier models are sparse; only the qualitative pattern (adequate at 1 percent, inadequate at 0.0001 percent) is claimed to be robust.
  • The distributional comparison only measured distance from ordinary user conversation data (LMSYS-Chat-1M) and explicitly does not measure distance to real adversarial attack traffic.
  • Both theorems assume approximately independent trials; adaptive red-teaming where a model detects it is being tested could change the underlying statistics, and the paper leaves open how much this would shift the boundary.
  • The illustrative numbers in Table 2 are meant to show structure, not to represent measured rates from any actual published evaluation procedure.

Why it matters

This gives AI labs, auditors, and regulators a way to check whether a 'no harmful outputs found' claim in a safety report or system card is actually statistically meaningful, rather than just assuming bigger test sets always help. It also tells the field when to stop asking for larger benchmarks and instead demand different kinds of evidence for extremely rare harm categories.

Terms in this paper

  • evidential ceiling · the maximum factor by which a single test result can shift belief about safety, given a fixed testing budget
  • null result · a test outcome where none of the prompts produced a harmful response
  • likelihood ratio (Λ) · a number showing how much better one hypothesis (unsafe vs. safe) explains the observed result than the other
  • Clopper-Pearson interval · a precise statistical method for estimating the plausible range of a true rate from an observed count of events
  • design effect · a correction factor accounting for how test prompts drawn from similar templates are not fully independent, shrinking the effective sample size

Original abstract (English)

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated e

Authors · Bandana Kaur

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Bandana Kaur et al., arXiv:2607.21735, CC BY 4.0