Go1 v0 -> v1
fell: 60 safer, 0 worse (aalen_johansen + fisher_bh, BH q<0.05 over 400 tests, paired)
Go1 v0 -> v1
fell: 60 safer, 0 worse
(aalen_johansen + fisher_bh, BH q<0.05 over 400 tests, paired)
fell
fell: 60 safer, 0 worse
-75.0
-81.2
-71.9
-71.9
-50.0
-37.5
-31.2
-43.8
-25.0
-9.4
-6.2
-6.2
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-34.4
-50.0
-46.9
-59.4
-37.5
-81.2
-75.0
-56.2
-43.8
-34.4
-6.2
-6.2
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-6.2
-6.2
-3.1
-18.8
-25.0
-59.4
-68.8
-43.8
-59.4
-9.4
-6.2
-9.4
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-9.4
-12.5
-40.6
-43.8
-81.2
-53.1
-21.9
-9.4
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
3.1
0.0
0.0
-9.4
-25.0
-50.0
-65.6
-43.8
-50.0
-21.9
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-31.2
-43.8
-50.0
-81.2
-21.9
-21.9
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-28.1
-40.6
-65.6
-53.1
-31.2
-12.5
-12.5
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-6.2
-12.5
-18.8
-31.2
-37.5
-40.6
-9.4
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-9.4
-18.8
-53.1
-31.2
-37.5
-18.8
-12.5
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-6.2
-15.6
-18.8
-43.8
-43.8
-28.1
-21.9
6.2
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-6.2
-28.1
-50.0
-40.6
-56.2
-31.2
-9.4
-6.2
0.0
0.0
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-25.0
-31.2
-71.9
-21.9
-25.0
-12.5
0.0
-3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-6.2
-25.0
-40.6
-50.0
-15.6
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-18.8
-15.6
-21.9
-25.0
-21.9
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-9.4
-25.0
-40.6
-28.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-9.4
-12.5
-21.9
-6.2
-6.2
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
-3.1
-21.9
-18.8
-31.2
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-6.2
-18.8
-28.1
-21.9
3.1
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
-3.1
12.5
-21.9
-9.4
-6.2
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
6.2
-3.1
-25.0
-34.4
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
160
170
180
190
200
0.05
0.1
0.15
0.2
0.25
0.3
0.35
0.4
0.45
0.5
0.55
0.6
0.65
0.7
0.75
0.8
0.85
0.9
0.95
1
push_pct_bw (rows: mu)
-85.0
0.0
+85.0
both 0% / both 100%
delta = rate_b - rate_a in pp; event=fell; negative is safer only if the event is a failure; corner square = significant; blank = no data
Worst-powered unchanged cell (push_pct_bw=100, mu=0.95, n=32/32) can only detect a 34pp rise. "0 worse" is an honest claim only above that.
Worst-powered cell overall: 34pp (push_pct_bw=10, mu=0.1).
fell: 60 safer, 0 worse
push_pct_bw mu n (a/b) rate_a (pp) rate_b (pp) delta (pp) p q verdict MDE 90 0.95 32/32 3.1 15.6 12.5 0.196 0.777 unchanged 25pp 120 0.5 32/32 93.8 100.0 6.2 0.492 1 unchanged 6pp 80 1 32/32 0.0 6.2 6.2 0.492 1 unchanged 23pp 120 0.9 32/32 90.6 93.8 3.1 1 1 unchanged 9pp 20 0.25 32/32 0.0 3.1 3.1 1 1 unchanged 23pp 130 0.7 32/32 100.0 100.0 0.0 1 1 unchanged 2pp 180 0.4 32/32 100.0 100.0 0.0 1 1 unchanged 2pp 150 0.85 32/32 100.0 100.0 0.0 1 1 unchanged 2pp 170 0.1 32/32 100.0 100.0 0.0 1 1 unchanged 2pp 160 0.1 32/32 100.0 100.0 0.0 1 1 unchanged 2pp
Worst-powered unchanged cell (push_pct_bw=100, mu=0.95, n=32/32) can only detect a 34pp rise. "0 worse" is an honest claim only above that. Worst-powered cell overall: 34pp (push_pct_bw=10, mu=0.1).
Methods
method pairing safer worse unchanged max MDE q fisher_bh · fell paired 60 0 340 34pp 0.05 mcnemar_bh · fell paired 55 0 345 34pp 0.05 logrank_bh (cause-specific) · fell paired 88 0 312 34pp 0.05
This is empirical evidence under a stated test envelope, not formal verification. Cells outside the envelope were not tested; cells inside it were tested at the stated replicate count and can only detect effects at or above the reported MDE.
Provenance
events: fell inferred: false non event: none n tests: 400 method: fisher_bh paired: true q: 0.05 t unit: seconds horizon: 5 n not comparable: 0 dropped a: 0 dropped b: 0 truncated: 0 n absent: 0 proofload version: 0.0.1 run id a: go1-v0 run id b: go1-v1 run salt: go1-g4-32:PRNGKey(1000000+rollout)