Poisson Labs Log in Contact

Poisson Labs · Direct lab audits

Two real receipts, unedited, and the audit that produces one.

An audit is the lab’s method on your system: we run it across the conditions that matter, hand you a fix list, and rerun it after the fix to see what moved. For teams facing a launch, a customer pilot or a change they cannot take back; run by the lab that builds the tools, a small number at a time.

Ask about an audit See what you get

How an audit runs

  1. 1

    Map

    We run your system across the conditions that matter and show you where it fails.

  2. 2

    Fix list

    A short list of changes, each tied to the failures it addresses.

  3. 3

    Re-test

    Same conditions, same seeds, after your change: what got better, what got worse.

  4. 4

    Receipt

    A written report and the runs behind it. Yours to keep, and to run again.

What we test

Agentic workflows
Tool-using and multi-step agents. Tool errors, slow responses, truncated outputs and retries become test conditions, and the report shows where in the workflow runs are lost.
Robot and control policies
Learned controllers, in your simulator. Friction, pushes, payloads and delays as axes; which checkpoint to ship; whether the retrain moved the boundary or only moved the failures.
Post-training runs Planned
Fine-tunes and reinforcement learning runs. Whether the new checkpoint is better across the conditions you care about or only on average, and where the reward was collected without the task being done.
Data pipelines Planned
Model steps inside a pipeline: extraction, labelling, routing. How much the output moves when nothing changed, and which inputs a new prompt or a new model made worse.

Agents and robot policies each have a published study behind them: the same run, twice and the Go1 failure boundary. Post-training runs and data pipelines are planned. We have not run either yet, so there is no receipt to show for them. If one of them is your problem, write to us anyway: it moves up.

What you keepExactly as the tool wrote them.

The report, the receipt, and every run behind them. The finding is stated with its limit: what moved, and what was too small for the run to detect.

These two are from our own studies, exactly as the tool wrote them. One is a fix that worked. The other is a fix that made things worse, which is the one you would want to know about before shipping. An audit’s written report sits on top of artifacts like these.

The fix worked. Go1 walking policy, retrained on the two gaps the map exposed: 60 of 400 conditions significantly safer, 0 worse. Open the receipt · report.md · the study
The fix made it worse. Microduck walking policy, a retrain meant to widen the boundary: 0 of 14 conditions safer, 14 worse. Open the receipt · report.md · the study

Empirical evidence under a stated test envelope, not formal verification. Conditions outside the envelope were not tested; conditions inside it can only resolve effects at or above the reported MDE.

Where this sits

Every study on this site is this method, applied to a system we chose. An audit is the same method, applied to yours.

The tools are what the work runs on, and the methods are written down. What is sold and what stays open is stated once, on the Sangfroid page.

Ask about an audit