4080120160200 0.20.40.60.81.0
PUSH, % BODYWEIGHT → ↑ FRICTION μ 32 TRIALS PER CELL PER VERSION
Go1 locomotion, baseline → retrained. Hover any cell.

A Go1 locomotion policy across 400 conditions of floor friction and lateral push. The map exposed two conditions training had never covered: pushes, and friction below 0.4. We retrained on exactly those two and ran the same 12,800 rollouts again: 60 conditions significantly safer, 0 worse.

Read the study →

Fig. 1Go1 locomotion, baseline against retrained: 400 conditions of floor friction (up) and lateral push in percent of bodyweight (across), 32 trials per cell, paired on seed. Green: the retrained policy fell significantly less often. Grey: no significant change; darker grey, a higher baseline fall rate. The worst-powered unchanged cell could only have detected a 34pp rise.

Research

  1. 2026.09 Where a 14-DoF Biped Falls Two checkpoints of a Microduck walking policy under a lateral push: a cliff and a slope, with a crossover between 0.600 and 0.625 m/s. A retrain meant to widen the boundary was worse in 14 of 14 cells. Microducklocomotion read reportreceiptretrain report
  2. 2026.08 Map the Failure Boundary A Go1 locomotion policy across 400 conditions of floor friction and lateral push. Retrained on the two gaps the map exposed: 60 conditions significantly safer, none worse. Go1locomotion read reportreceipt
  3. 2026.05 Monte Found a Decorative Channel and a Reward Exploit in One of Our Benchmarks A communication channel in a MARL benchmark that looked load-bearing and was not, and a reward function that let trivial policies win. MARLcommunication read
  4. 2026.05 Where V-JEPA 2.1's Dense Features Hold Up (and Where They Don't) Across four model sizes, dense features predict failure under temporal corruption and not under image noise, and the 2B model is less robust than the 1B on three of five perturbations. V-JEPA 2.1representation read code
  5. 2026.01 10,924x: The Instability Bomb at 1.7B Scale Scaled the reproduction from 10M to 1.7B parameters: unconstrained Hyper-Connections reached 10,924× signal amplification; the Sinkhorn-constrained version held at 1.0. mHCscaling read
  6. 2026.01 DeepSeek's mHC: When Residual Connections Explode Reproduced the 9× signal amplification in DeepSeek's Hyper-Connections at 10M parameters, and the fix: Sinkhorn-Knopp normalisation of the mixing matrices. mHCresiduals read

Instrument

Proofload

Find where a policy or an agent breaks, and prove whether a change moved it.

Name the conditions. Proofload runs your system across them, before and after a change, and returns what moved, and what was too small for the run to detect. Your physics, your agent, your outcome. It brings the sweep, the seeds, the store and the statistics.

Open source, releasing soon. About the instrument, or write to be told when it ships.

0.550.5750.60.6250.650.6750.7 0.61.0
LATERAL PUSH, m/s → ↑ FRICTION μ 100 TRIALS PER CELL PER VERSION
Microduck walking, checkpoint 2250 → checkpoint 5999. Hover any cell.
Fig. 2What the instrument returns. A Microduck walking policy at two checkpoints, checkpoint 2250 against checkpoint 5999, across 14 conditions of lateral push and floor friction, 100 trials per cell, paired on seed. Green: the later checkpoint fell significantly less often; red, more; grey, no significant change, where the run could only have detected a 8pp difference. The ranking flips between 0.600 and 0.625 m/s. Runs 20260909T010313Z-5abe and 20260909T010448Z-a6f1. report.md

About

A small lab, working in the open, on where learned systems stop working and how to tell whether a fix helped.

We build the instruments, use them on our own work, and publish the results with what they could not detect. We take on a small amount of outside work close to the research: robustness testing, evaluation design, RL systems.

work@poissonlabs.ai