Writing 5 min read

Judge the Application, Not 'AI': A Rubric for Clinical AI

Why I judge each clinical AI application on its own evidence, not the category, and the rubric I run before I trust one.

Judge the Application, Not 'AI': A Rubric for Clinical AI

Photo by Pavel Danilyuk on Pexels

The short answer

The question is AI good for medicine is the wrong one, and in my experience it slows down the tools that work while giving cover to the ones that do not. AI is not one thing. An ambient scribe and a sepsis predictor share almost nothing that matters. I judge each application on its own evidence, against a rubric, and I let the evidence, not the category, decide.

Every few weeks someone asks me whether AI is good for medicine, and I have stopped answering the question, because it is the wrong one. As one STAT piece put it, asking whether AI improves health care is like asking whether lasers improve surgery. It depends entirely on the laser, the procedure, and the surgeon. AI is not a single force. An ambient scribe that drafts a note and a model that predicts sepsis share the letters A and I and almost nothing else that matters.

So here is how I actually decide. I never ask if AI is good for medicine. I ask if this specific application, in this specific setting, has earned my trust with evidence. I judge each one on two things: how strong the evidence is that it works, and how much damage it does when it is wrong. The category tells me nothing. The evidence tells me everything.

AI is not one thing

The most useful thing you can do with any clinical AI is place it on two axes: how good the evidence is, and how bad the failure is. I map every application before I form an opinion. On one axis, how strong is the evidence that it actually works in the real world, not on the vendor's internal benchmark. On the other, what happens when it fails: a mild annoyance, or a missed diagnosis. An ambient scribe with good evidence and a low failure cost sits in a very different place from an autonomous triage model with thin evidence and a high failure cost. Same three letters. Opposite decisions.

Judge each clinical AI on evidence and on the cost of being wrong

Do not deployhigh stakes, thin evidenceDeploy with oversighthigh stakes, strong evidencePilot and watchlow stakes, thin evidenceShip itlow stakes, strong evidenceEvidence strength →Consequence of error →

The same category, opposite evidence

Two tools can both be called predictive clinical AI and be worlds apart on evidence, and the sepsis story is the one I keep coming back to. A widely deployed proprietary sepsis prediction model was marketed with an internal accuracy, an AUC around 0.76 to 0.83. When independent researchers validated it on tens of thousands of real hospitalizations, published in JAMA Internal Medicine, it scored 0.63. It missed 67 percent of sepsis cases while firing alerts on 18 percent of all patients, roughly one useful alert for every eight it raised. That is not a rounding error, it is the difference between a tool you trust and a tool that buries clinicians in false alarms. Meanwhile ambient documentation tools, a completely different application, have gathered far more convincing evidence of giving clinicians time back. Same word, AI. Completely different verdict. The category did not tell me which was which. The evidence did.

Vendor-reported AUC (internal)
0.76-0.83
Independent validation (JAMA Intern Med)
0.63

The rubric I use

Before I trust any clinical AI, I run it through the same short rubric, and if it cannot answer these, the answer is no. Is there evidence it works on people like mine, not just the vendor's cohort? Was it validated by someone without a stake in the result? Do I know what happens when it is wrong, and who catches it? Does it fit the workflow, or fight it? And is the benefit worth the alerts it adds? Score it honestly. A high score earns a supervised deployment. A low score earns a pilot, or a polite no. Run the five questions below on whatever tool is in front of you.

Should you trust this clinical AI?

Five questions before you deploy any application.

Is there evidence it works on a population like yours, not just the vendor's?

Was it validated independently, by someone without a stake in the result?

Do you know exactly what happens when it is wrong, and who catches it?

Does it fit the clinical workflow rather than fight it?

Is the benefit worth the alert burden it adds?

0%
Answer all five

None of this is anti-AI. I am for the tools that earn it. But is AI good for medicine is a question that lets a bad sepsis model hide behind a good scribe, and a good scribe get tarred by a bad sepsis model. Drop the category. Judge the application. Ask for the evidence, and be willing to walk away when it is not there.

Key takeaways
  • Is AI good for medicine is the wrong question; judge each application on its own evidence.
  • Map every clinical AI on two axes: strength of evidence, and consequence of being wrong.
  • A widely used sepsis model scored 0.63 on independent validation versus a vendor-reported 0.76 to 0.83, missing 67 percent of cases.
  • Independent validation on a population like yours matters more than any vendor benchmark.
  • Weigh the benefit against the alert burden; a tool that buries clinicians in false alarms is a failure.
  • Run a short rubric before deploying, and be willing to say no.

Frequently asked

Why is is AI good for medicine the wrong question?

Because AI is not one thing. An ambient scribe and a sepsis predictor are different applications with different evidence, so a single verdict hides the truth about both.

What is the Epic Sepsis Model example?

A widely deployed proprietary sepsis model that scored an AUC of 0.63 in an independent JAMA Internal Medicine validation, far below vendor-reported figures, and missed 67 percent of sepsis cases.

How should I evaluate a clinical AI tool?

On evidence it works for a population like yours, independent validation, a clear failure and oversight plan, workflow fit, and whether the benefit beats the alert burden.

Why does independent validation matter?

Because vendor benchmarks are often measured on favorable internal cohorts. Independent validation on real, local populations is what predicts real-world performance.

Does this mean I should avoid clinical AI?

No. It means back the applications with strong evidence and a clear oversight plan, and decline the ones without, rather than judging the whole category at once.

Sources

Naveen Kumar

Naveen Kumar

Healthcare engineering and product executive in Pittsburgh. 15+ years building AI-first patient access, a decade at Treatspace.

Read next