Writing 5 min read

Never a Meat Proxy: Human-in-the-Loop Discipline for Clinical AI

Why no AI output should reach a patient unverified, and the principles I use to keep the human in the loop genuinely in the loop.

Never a Meat Proxy: Human-in-the-Loop Discipline for Clinical AI

Photo by RDNE Stock project on Pexels

The short answer

The rule I hold for clinical AI is simple: no AI output reaches a patient or a chart until a responsible human has read it, understood it, and verified it. Simon Willison calls the alternative being a meat proxy, a person who passively relays what the model said. In medicine that is not a style problem, it is a safety failure, because the model is confidently wrong often enough to matter and automation bias makes us wave it through.

Every clinical AI deployment eventually comes down to one question: who checks the output before it matters? The demo never shows this part. It shows the model drafting a note or suggesting an order in seconds. What it does not show is the moment a tired clinician clicks approve without reading, and the model's confident mistake becomes a fact in someone's chart. That moment is where clinical AI is actually safe or not.

My rule is simple and I do not bend it. No AI output reaches a patient or a chart until a responsible human has read it, understood it, and verified it. Simon Willison has a sharp name for the alternative, borrowed from Niklas Gruhn: a meat proxy, a person who passively relays what the model said. In medicine that is not a style problem, it is a safety failure. Here are the principles I use to keep the human in the loop actually in the loop.

Principle 01

Verification is the job, not a formality.

Willison, borrowing the term from Niklas Gruhn, describes the meat proxy: someone who takes AI output and relays it without reading, understanding, or validating it. His fix is to actually engage, read it, understand it, validate it, then put it in your own words as a certificate that you did. In clinical work that certificate is the whole point. A clinician who signs an AI-drafted note without reading it has not supervised the AI. They have laundered its output through a human name.

Principle 02

Automation bias is real, and it is measurable.

This is not hypothetical caution. In a study of trained experts making AI-assisted decisions under time pressure, researchers measured a 7 percent automation bias rate, where initially correct human judgments were overturned by erroneous AI advice. Seven percent sounds small until you multiply it across every AI-assisted decision in a health system. And pressure made it worse. The lesson is not to distrust every model, it is to design for the fact that a plausible, confident, wrong answer is exactly the one we are most likely to accept.

Principle 03

Put the human where the harm is.

You cannot deeply verify everything, and pretending you will is how verification quietly becomes rubber-stamping. So I put the strongest human review where an error reaches a patient or a permanent record, and I let low-stakes outputs run lighter. A drafted internal summary and an order that changes treatment do not deserve the same scrutiny. Name the high-harm outputs explicitly, and make the human step there non-negotiable.

Spend human review where an error can reach a patient

Wasted effortdeep review, low stakesRightdeep review where it mattersFine as islight touch, low stakesDanger zonerubber-stamped, high stakesHarm potential →Review depth →

A clinician who signs an AI note without reading it has not supervised the AI. They have laundered its output through a human name.

Naveen Kumar
Principle 04

Design against the rubber stamp.

The failure I watch for most is the review step that exists on paper but not in practice. If the interface makes approving as effortless as breathing, people will approve on autopilot, especially when the model is usually right, which is what makes the occasional wrong one so dangerous. So I design friction where it matters: show the evidence, require the reviewer to engage with the specific claim, make them do something that only makes sense if they actually read it. The point is not to slow people down for its own sake. It is to make real review the path of least resistance.

Principle 05

A signature means responsibility, so keep it human.

In the end, a person puts their name on what reaches the patient, and that name has to mean something. The model can draft, suggest, and speed things up. It cannot be accountable. So the human in the loop is not there as a courtesy, they are there as the accountable party, and everything about the system should reinforce that they are responsible for what they approve. When the name on the chart is a human's, the verification behind it has to be too.

Human-in-the-loop is not a checkbox on a compliance form, it is the load-bearing safety mechanism in almost every clinical AI system I would trust. And it only works if the human is really there: reading, understanding, verifying, and willing to say no. Build the system so that is the easy path, measure whether it is actually happening, and never let the human become a rubber stamp with a pulse. The model can do a great deal. Owning the outcome is still ours.

Key takeaways
  • No AI output should reach a patient or chart until a responsible human has read, understood, and verified it.
  • Passively relaying model output, being a meat proxy, adds no safety and keeps the blame.
  • Automation bias is measurable: one study found 7 percent of correct human judgments overturned by wrong AI advice, worse under time pressure.
  • Spend deep review where an error can reach a patient, and go lighter on low-stakes output.
  • Design against the rubber stamp: make real engagement the path of least resistance.
  • A signature means accountability, which a model cannot hold, so keep the responsible human genuinely responsible.

Frequently asked

What does human-in-the-loop mean for clinical AI?

That a responsible person reads, understands, and verifies AI output before it affects a patient or the record, rather than passing it along unchecked.

What is a meat proxy?

Simon Willison's term, via Niklas Gruhn, for someone who relays AI output without reading, understanding, or validating it, adding no value and no safety.

What is automation bias?

The tendency to over-trust automated or AI recommendations, including accepting a wrong one over a correct human judgment. It has been measured around 7 percent in AI-assisted expert decisions and worsens under time pressure.

Do you have to verify every AI output equally?

No. Concentrate deep review where an error can reach a patient or a permanent record, and apply a lighter touch to low-stakes output.

How do you stop review from becoming a rubber stamp?

Design friction where it matters: show the evidence, require engagement with the specific claim, and measure whether real review is happening rather than reflexive approval.

Sources

Naveen Kumar

Naveen Kumar

Healthcare engineering and product executive in Pittsburgh. 15+ years building AI-first patient access, a decade at Treatspace.

Read next