Go1 v0 -> v1

fell: 60 safer, 0 worse
(aalen_johansen + fisher_bh, BH q<0.05 over 400 tests, paired)

Go1 v0 -> v1 fell: 60 safer, 0 worse (aalen_johansen + fisher_bh, BH q<0.05 over 400 tests, paired) fell fell: 60 safer, 0 worse -75.0 -81.2 -71.9 -71.9 -50.0 -37.5 -31.2 -43.8 -25.0 -9.4 -6.2 -6.2 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -34.4 -50.0 -46.9 -59.4 -37.5 -81.2 -75.0 -56.2 -43.8 -34.4 -6.2 -6.2 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -6.2 -6.2 -3.1 -18.8 -25.0 -59.4 -68.8 -43.8 -59.4 -9.4 -6.2 -9.4 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -9.4 -12.5 -40.6 -43.8 -81.2 -53.1 -21.9 -9.4 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 3.1 0.0 0.0 -9.4 -25.0 -50.0 -65.6 -43.8 -50.0 -21.9 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -31.2 -43.8 -50.0 -81.2 -21.9 -21.9 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -28.1 -40.6 -65.6 -53.1 -31.2 -12.5 -12.5 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -6.2 -12.5 -18.8 -31.2 -37.5 -40.6 -9.4 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -9.4 -18.8 -53.1 -31.2 -37.5 -18.8 -12.5 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -6.2 -15.6 -18.8 -43.8 -43.8 -28.1 -21.9 6.2 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -6.2 -28.1 -50.0 -40.6 -56.2 -31.2 -9.4 -6.2 0.0 0.0 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -25.0 -31.2 -71.9 -21.9 -25.0 -12.5 0.0 -3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -6.2 -25.0 -40.6 -50.0 -15.6 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -18.8 -15.6 -21.9 -25.0 -21.9 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -9.4 -25.0 -40.6 -28.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -9.4 -12.5 -21.9 -6.2 -6.2 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 -3.1 -21.9 -18.8 -31.2 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -6.2 -18.8 -28.1 -21.9 3.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 -3.1 12.5 -21.9 -9.4 -6.2 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 6.2 -3.1 -25.0 -34.4 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 200 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95 1 push_pct_bw (rows: mu) -85.0 0.0 +85.0 both 0% / both 100% delta = rate_b - rate_a in pp; event=fell; negative is safer only if the event is a failure; corner square = significant; blank = no data Worst-powered unchanged cell (push_pct_bw=100, mu=0.95, n=32/32) can only detect a 34pp rise. "0 worse" is an honest claim only above that. Worst-powered cell overall: 34pp (push_pct_bw=10, mu=0.1).

fell: 60 safer, 0 worse

push_pct_bwmun (a/b)rate_a (pp)rate_b (pp)delta (pp)pqverdictMDE
900.9532/323.115.612.50.1960.777unchanged25pp
1200.532/3293.8100.06.20.4921unchanged6pp
80132/320.06.26.20.4921unchanged23pp
1200.932/3290.693.83.111unchanged9pp
200.2532/320.03.13.111unchanged23pp
1300.732/32100.0100.00.011unchanged2pp
1800.432/32100.0100.00.011unchanged2pp
1500.8532/32100.0100.00.011unchanged2pp
1700.132/32100.0100.00.011unchanged2pp
1600.132/32100.0100.00.011unchanged2pp

Worst-powered unchanged cell (push_pct_bw=100, mu=0.95, n=32/32) can only detect a 34pp rise. "0 worse" is an honest claim only above that.
Worst-powered cell overall: 34pp (push_pct_bw=10, mu=0.1).

Methods

methodpairingsaferworseunchangedmax MDEq
fisher_bh · fellpaired60034034pp0.05
mcnemar_bh · fellpaired55034534pp0.05
logrank_bh (cause-specific) · fellpaired88031234pp0.05
This is empirical evidence under a stated test envelope, not formal verification. Cells outside the envelope were not tested; cells inside it were tested at the stated replicate count and can only detect effects at or above the reported MDE.

Provenance