Johns Hopkins, one of the most technically sophisticated hospital systems on the planet, tested every leading frontier AI agent on real healthcare tasks. The results landed last week. Only 28 percent of complex tasks were completed on a first attempt. Fewer than 8 percent of agents remained consistently successful across repeated runs (Healthcare IT News, July 2026). These are not toy benchmarks. These are clinical workflows, documentation chains, and operational tasks that real staff perform every day.

The instinct is to read that number and conclude the agents are not ready. Wrong lesson. Johns Hopkins read those numbers and did something most organizations have not considered: they stopped asking which agent is best and started asking which tasks are simple enough.

The Capability Trap

Most companies shopping for AI agents are asking the wrong question. They compare models. They evaluate capability scores. They read vendor benchmarks that show agents completing elaborate multi-step processes in controlled environments. Then they deploy those agents against their actual workflows and watch them fail at rates that make the pilot look broken.

The pilots are fine. The scope is wrong.

A frontier AI agent that can reason through a twelve-step process in a demo environment will still fail when the twelfth step depends on a system it cannot access, a permission it does not have, or a context window that lost track of step four. The model works fine. The assumption that complex means capable is where it falls apart.

Johns Hopkins found this out with data instead of opinions. Twenty-eight percent on first attempt. That number did not improve by switching to a different model. It improved by changing what the agents were asked to do.

What Simplicity Looks Like in Practice

The teams inside Johns Hopkins that got consistent results did not deploy the smartest agent on the hardest task. They broke complex clinical workflows into discrete, well-defined steps. One agent handles one handoff. One agent processes one form. One agent checks one set of records against one standard.

That sounds underwhelming until you see the math. An agent that completes a simple, bounded task at 94 percent reliability, running across hundreds of cases a day, produces more value than an agent that attempts a complex multi-step process and fails 72 percent of the time. Reliability compounds. Ambition does not.

This is the pattern showing up everywhere, not just healthcare. The 97 percent of executives who have deployed agents but only 29 percent seeing ROI (Writer, July 2026) are largely stuck in the same trap. They deployed capable agents against complex work and mistook deployment for results.

The First Step Most Teams Skip

Before Johns Hopkins let a single agent touch a production workflow, they benchmarked. Not the way vendors benchmark, running controlled tests designed to show the agent at its best. They built evaluation frameworks that mirror the actual complexity, the actual data quality, and the actual failure modes their teams face daily.

That is the step almost nobody takes. Most organizations go straight from vendor demo to pilot deployment. The demo worked, so the pilot should work. When it does not, the conclusion is that AI is not ready, or that the team needs a better model. Usually neither. They need a better match between the agent’s capability and the task’s complexity.

If you run an organization that is deploying or evaluating AI agents, here is the honest assessment: the agent you are testing can probably do about a quarter of what your vendor showed you, at production quality, under real conditions. Every technology has always had a gap between demo and deployment. AI agents are no different. The question is whether you adjust the scope or keep throwing the same agent at the same complex task and hoping the next model release fixes it.

Why This Matters Beyond Healthcare

Johns Hopkins is not a typical enterprise. But the finding transfers directly. The 28 percent number will vary by industry, by workflow complexity, by data quality. The pattern will not. Complex multi-step agent deployments fail at rates that make ROI nearly impossible to prove. Simple, scoped agent deployments succeed at rates that make ROI obvious within weeks.

The companies getting real results from AI agents right now share one trait: they picked the simplest systems that solve real problems. An agent that reliably pulls data from three sources and formats a daily report is worth more than an agent that sometimes completes an entire quarterly analysis but fails unpredictably on the third step.

Simplicity is the deployment strategy that actually works. Johns Hopkins proved it with data. Everyone else is running the experiment on their own budget.