Agentic AI Is Coming to Medicine. Cage It Where It Counts.
Agentic AI just beat physicians on a controlled test. Real medicine is messier. Where I let the agent run, and where I cage it.

Photo by Pavel Danilyuk on Pexels
Agentic AI in medicine is real, and the latest results are genuinely impressive. My take is that impressive-in-a-clean-study and safe-in-a-messy-clinic are different things, and the gap between them is exactly where autonomy should be caged. I welcome the capability. I still give the agent one bounded room to work in, with a human on the door, until the evidence catches up to the demo.
Eric Topol had a piece recently, Agentic AI Comes to Medicine, that is worth reading in full. He walks through two agentic systems published in Nature: MIRA for the emergency department and AMIE for outpatient care. The headline numbers are striking. MIRA reached 87.8 percent diagnostic accuracy against 78.1 percent for board-certified physicians, and on appendicitis it hit 100 percent against 88 percent. If you build in this space, results like that get your attention, and they should.
But Topol is careful, and so am I. Those systems ran on text-only inputs, clean and complete data, and patient actors or capped interactions. Real medicine is none of those things. It is incomplete, conflicting, and messy, which is where confident systems get into trouble. So my stance is not to dismiss agentic AI, and not to hand it the clinic either. It is to welcome the capability and cage the autonomy, giving the agent room to work exactly where the stakes and the mess are both low.
The results are real
I want to be honest that agentic AI just cleared a bar narrow AI never did: end-to-end reasoning that beats physicians on a controlled test. For years, medical AI meant a narrow model doing one step, usually flagging a diagnosis for a human to act on. What Topol describes is different: systems that reason across the whole encounter and make decisions, and do it well on the test they were given. The chart below is the headline. On a controlled diagnostic evaluation, the agentic system beat board-certified physicians. That is a real milestone, and pretending otherwise would be dishonest.
Clean study, messy clinic
The distance between that number and a safe deployment is the same distance between a clean dataset and a real patient. Here is what the number does not include. Real patients arrive with incomplete and conflicting information, not the tidy vignettes these systems were tested on. They speak, they hedge, they leave things out. The evaluations used text only, and often patient actors, with the interaction capped. None of that is a knock on the research, it is the research being honest about its limits. But it means the gap between 87.8 percent on a controlled test and safe autonomous care in a real emergency department is enormous. I have watched too many tools ace the demo and stumble on the mess to confuse the two.
Where agentic AI fits, and where it does not
So I do not ask whether to use agentic AI. I ask where, and the answer falls out of two questions: how high are the stakes, and how messy is the input. Map any task on those two axes. Low stakes and clean input is where I let an agent run with light oversight. High stakes or messy input is where I cage it: a bounded step, a human on the decision, a deterministic backbone around it. The dangerous corner is high stakes and messy input, which unfortunately describes a lot of real medicine, and which is exactly where a confident agent will hurt someone. Autonomy is a dial, and I set it by the corner the task lives in.
I am optimistic about agentic AI in medicine, and I am in no hurry to hand it the keys. Both can be true. The research is real progress, and real medicine is still messier than any benchmark. So I welcome each new result, I read the caveats as carefully as the headline, and I keep the agent caged where the stakes and the mess are high. When the evidence moves from clean studies to messy clinics, I will move the dial. Not before.
- Agentic AI in medicine is real; recent systems beat physicians on controlled diagnostic tests.
- One agentic system reached 87.8 percent diagnostic accuracy versus 78.1 percent for board-certified physicians.
- Those results used text-only, clean, complete data and capped interactions, unlike real medicine.
- Decide where to use agentic AI by two axes: stakes and input messiness.
- Cage autonomy where stakes are high or input is messy; let it run where both are low.
- Welcome the capability, read the caveats, and move the autonomy dial only as evidence moves from clean studies to messy clinics.
Frequently asked
Is agentic AI good enough for medicine?
It is impressive on controlled tests, with one system beating physicians on diagnostic accuracy, but those tests used clean, text-only data unlike real, messy clinical encounters.
What did Eric Topol report about agentic AI?
Two Nature systems, MIRA and AMIE, with MIRA reaching 87.8 percent diagnostic accuracy versus 78.1 percent for physicians, alongside clear caveats about text-only, clean data.
Where should agentic AI be used in healthcare?
Where stakes and input messiness are both low, with heavier human oversight as either rises. High-stakes, messy tasks are not ready for autonomy.
What is the risk of agentic AI in medicine?
A confident system meeting incomplete, conflicting real-world data can be wrong with conviction, which is most dangerous in high-stakes, messy situations.
How is this different from earlier medical AI?
Earlier tools were narrow, doing one step for a human to act on. Agentic systems reason across the whole encounter and make decisions, which raises both the potential and the risk.