ACTIVE  ·  BUILDING  ·  v1.0 2026-09-15  ·  JL:IOTA:001
No. 154 · 2026-09-13

The Test Still Touches Production

DISPATCH  ·  LOGGED WITH MAI

OpenAI and Anthropic are now saying the quiet part in public: powerful agents need outside review close to the work.

Good.

The useful lesson is not that every company needs a frontier-lab safety program. Most do not. The lesson is smaller and more annoying.

If the agent can touch a real system, the review has to be inside the workflow.

Not after the fact. Not when the report is done. Not when the demo looks strange and someone finally asks what it changed.

Inside.

Before it starts, name the job. During the work, log the steps that matter. At the boundary, make it stop. Afterward, review both the output and the path it took to get there.

This sounds boring because control usually does.

But the OpenAI RubyGems reporting is the warning shot. An agent does not need bad intent to create a real mess. It can misunderstand the task, overreach, publish into a place nobody counted as live, or complete the wrong thing fast enough that the team notices too late.

The word test does not shrink the blast radius.

So stop treating review like polish at the end of the run.

Give the agent a scope. Give it a reviewer. Give it a log. Give it a reason to stop.

Then run the test.

LOGGED WITH MAI  ·  2026-09-13  ·  No. 154
← All Dispatches