A Go1 locomotion policy across 400 conditions of floor friction and lateral push. The map exposed two conditions training had never covered: pushes, and friction below 0.4. We retrained on exactly those two and ran the same 12,800 rollouts again: 60 conditions significantly safer, 0 worse.
Read the study →Fig. 1Go1 locomotion, baseline against retrained: 400 conditions of floor friction (up) and lateral push in percent of bodyweight (across), 32 trials per cell, paired on seed. Green: the retrained policy fell significantly less often. Grey: no significant change; darker grey, a higher baseline fall rate. The worst-powered unchanged cell could only have detected a 34pp rise.
Research
- 2026.09 Where a 14-DoF Biped Falls Two checkpoints of a Microduck walking policy under a lateral push: a cliff and a slope, with a crossover between 0.600 and 0.625 m/s. A retrain meant to widen the boundary was worse in 14 of 14 cells.
- 2026.08 Map the Failure Boundary A Go1 locomotion policy across 400 conditions of floor friction and lateral push. Retrained on the two gaps the map exposed: 60 conditions significantly safer, none worse.
- 2026.05 Monte Found a Decorative Channel and a Reward Exploit in One of Our Benchmarks A communication channel in a MARL benchmark that looked load-bearing and was not, and a reward function that let trivial policies win.
- 2026.05 Where V-JEPA 2.1's Dense Features Hold Up (and Where They Don't) Across four model sizes, dense features predict failure under temporal corruption and not under image noise, and the 2B model is less robust than the 1B on three of five perturbations.
- 2026.01 10,924x: The Instability Bomb at 1.7B Scale Scaled the reproduction from 10M to 1.7B parameters: unconstrained Hyper-Connections reached 10,924× signal amplification; the Sinkhorn-constrained version held at 1.0.
- 2026.01 DeepSeek's mHC: When Residual Connections Explode Reproduced the 9× signal amplification in DeepSeek's Hyper-Connections at 10M parameters, and the fix: Sinkhorn-Knopp normalisation of the mixing matrices.
Instrument
Proofload
Find where a policy or an agent breaks, and prove whether a change moved it.
Name the conditions. Proofload runs your system across them, before and after a change, and returns what moved, and what was too small for the run to detect. Your physics, your agent, your outcome. It brings the sweep, the seeds, the store and the statistics.
Open source, releasing soon. About the instrument, or write to be told when it ships.
About
A small lab, working in the open, on where learned systems stop working and how to tell whether a fix helped.
We build the instruments, use them on our own work, and publish the results with what they could not detect. We take on a small amount of outside work close to the research: robustness testing, evaluation design, RL systems.