Twenty-Eight Percent
Johns Hopkins ran frontier AI agents against real clinical workflows last week. Not vendor demos. Not controlled benchmarks. Actual documentation chains, operational tasks, the work staff do every day. Twenty-eight percent of complex tasks completed on a first attempt. Fewer than 8 percent of agents stayed consistent across repeated runs.
The instinct is to blame the models. Johns Hopkins didn’t. They changed the scope instead. Teams that broke multi-step processes into single, bounded tasks — one agent, one handoff, one check — got completion rates above 90 percent. Not by upgrading the model. By downgrading the ask.
That math matters. An agent completing a simple task at 94 percent reliability across hundreds of daily cases produces more than one attempting a twelve-step process and failing seven times out of ten. Reliability compounds. Ambition doesn’t.
This tracks with a broader pattern. Ninety-seven percent of executives have deployed agents. Twenty-nine percent report ROI. The gap isn’t capability. It’s scope. Most teams went from vendor demo to production and mistook deployment for results.
The fix isn’t waiting for the next model release. It’s asking a smaller question. One that the agent can actually answer, every time, at production quality. Then asking the next one.