20260905T185407Z-0313 vs 20260905T195009Z-18ac: failed_assertion

failed_assertion: 0 safer, 1 worse
(aalen_johansen + mcnemar_bh, BH q<0.05 over 328 tests, paired)

Omnibus

"no cell changed" is rejected (p = 0.000249938 < q 0.05) over 328 cells x 1 event

min-p of location and dispersion; 4000 within-cell permutations, pair_swap, seed 0

p floor 2.50e-04: the smallest p 4000 permutations can express; a p at the floor is a bound, not a measurement

eventmean delta (pp)prms delta (pp)p
failed_assertion0.40.01755.42.50e-04

informative cells: 56 of 328; 272 unchanged by every exchange

Fault lines

grouping temperature (axis), 2 groups; cells per group min 164 / median 164 / max 164; 0 degenerate; 2 of 2 groups resolve 20pp at alpha 0.025

refused task_id (axis): underpowered; 0 of 164 groups can resolve a change of 20.0 points in a group's event rate (group MDE 0.3615 at 64 and 64 units per arm, alpha 0.000305 = q/164), and you accepted 20.0 points (--max-group-mde 0.2000). --max-group-mde 0.3616 admits this grouping as it stands. next: axis

"no cell in the group changed" tested by BH across 2 groups at q 0.05: 2 rejected

group p floor 2.50e-04: the same 4000 permutations; a group p at the floor is a bound, not a measurement

temperaturecellsinformativeunchanged by every exchangeunits (a/b)group MDE (pp)discordant pairspqverdict
0.0164141505248/52483.0pp1380.02250.0225rejected
0.7164421225248/52483.0pp2192.50e-045.00e-04rejected

Cells

"this cell did not change" tested by BH within each of 2 rejecting groups at q 0.05: 1 of 328 rejected

MDE is ADR-018's smallest detectable increase from arm A's rate, capped at the headroom, so on a cell at ceiling it reads the headroom and is not a power number.

temperature=0.0: 164 tests, 0 rejected; cell MDE median 23pp / max 34pp

temperature=0.7: 164 tests, 1 rejected; cell MDE median 23pp / max 34pp

celleventdelta (pp)qMDE
task_id=HumanEval/129, temperature=0.7failed_assertion62.53.13e-0434pp

Bound

depth 3 (3 configured); outer-nodes false discovery rate at most 0.432 = 2 x 3 x 0.05 x 1.44

Yekutieli (2008) Corollary A.2 assumes each p-value is independent of its ancestors; this tree's group and root statistics aggregate their own cells, so that assumption is not established here.

20260905T185407Z-0313 vs 20260905T195009Z-18ac: failed_assertion failed_assertion: 0 safer, 1 worse (aalen_johansen + mcnemar_bh, BH q<0.05 over 328 tests, paired) failed_assertion failed_assertion: 0 safer, 1 worse 0 0.7 HumanEval/0 HumanEval/103 HumanEval/109 HumanEval/114 HumanEval/12 HumanEval/125 HumanEval/130 HumanEval/136 HumanEval/141 HumanEval/147 HumanEval/152 HumanEval/158 HumanEval/163 HumanEval/21 HumanEval/27 HumanEval/32 HumanEval/38 HumanEval/43 HumanEval/49 HumanEval/54 HumanEval/6 HumanEval/65 HumanEval/70 HumanEval/76 HumanEval/81 HumanEval/87 HumanEval/92 HumanEval/98 temperature (rows: task_id) -65.0 0.0 +65.0 both 0% / both 100% delta = rate_b - rate_a in pp; event=failed_assertion; negative is safer only if the event is a failure; corner square = significant; blank = no data Worst-powered unchanged cell (temperature=0.7, task_id=HumanEval/110, n=32/32) can only detect a 34pp rise. "1 worse" is an honest claim only above that. Worst-powered cell overall: 34pp (temperature=0.7, task_id=HumanEval/110).

failed_assertion: 0 safer, 1 worse

temperaturetask_idn (a/b)rate_a (pp)rate_b (pp)delta (pp)pqverdictMDEinterval (pp)
0.7HumanEval/12932/3237.5100.062.51.91e-066.26e-04worse34pp[12.2, 85.8]
0HumanEval/12932/3262.5100.037.54.88e-040.0801unchanged29pp[-5.1, 68.6]
0.7HumanEval/3232/3262.584.421.90.09231unchanged29pp[-23.7, 58.9]
0HumanEval/14232/3234.453.118.80.181unchanged34pp[-27.6, 57.4]
0HumanEval/10032/3225.040.615.60.3021unchanged34pp[-31.4, 55.8]
0.7HumanEval/9132/3287.5100.012.50.1251unchanged12pp[-22.3, 45.9]
0HumanEval/8332/3284.496.912.50.2191unchanged16pp[-24.9, 47.3]
0.7HumanEval/13432/3268.881.212.50.1251unchanged26pp[-22.3, 45.9]
0.7HumanEval/10032/3231.243.812.50.2191unchanged34pp[-24.9, 47.3]
0.7HumanEval/9332/3281.290.69.40.251unchanged19pp[-24.5, 42.5]
0.7HumanEval/8332/3262.571.99.40.5081unchanged29pp[-31.2, 47.1]
0HumanEval/8932/3225.034.49.40.6071unchanged34pp[-36.7, 51.4]
0HumanEval/532/3293.8100.06.20.51unchanged6pp[-26.6, 38.9]
0HumanEval/9132/3290.696.96.20.6251unchanged9pp[-29.0, 40.6]
0.7HumanEval/1032/3271.978.16.20.6251unchanged25pp[-29.0, 40.6]
0.7HumanEval/10232/326.212.56.20.51unchanged28pp[-26.6, 38.9]
0.7HumanEval/15932/329.415.66.20.6251unchanged29pp[-29.0, 40.6]
0.7HumanEval/11932/3215.621.96.20.7271unchanged32pp[-33.1, 43.8]
0.7HumanEval/10832/3243.850.06.20.6251unchanged34pp[-29.0, 40.6]
0.7HumanEval/8932/3237.543.86.20.7541unchanged34pp[-35.0, 45.4]

Worst-powered unchanged cell (temperature=0.7, task_id=HumanEval/110, n=32/32) can only detect a 34pp rise. "1 worse" is an honest claim only above that.
Worst-powered cell overall: 34pp (temperature=0.7, task_id=HumanEval/110).

Intervals

intervals: 99.9848% (alpha* = 0.000152439 = q x 1 / 328 over the flat BH family), basis score_rd_paired

interval excludes zero, verdict unchanged: 0 cells

verdict significant, interval covers zero: 0 cells

This is empirical evidence under a stated test envelope, not formal verification. Cells outside the envelope were not tested; cells inside it were tested at the stated replicate count and can only detect effects at or above the reported MDE.

Provenance