Writing 5 min read

The FDA Is Coming for Generative AI: What Validated Will Mean

The FDA released a discussion paper on generative AI devices. It quietly ends the era of regulating AI as a locked box, and it changes what validated means.

The FDA Is Coming for Generative AI: What Validated Will Mean

Photo by Atharva Sune on Pexels

The short answer

The FDA released a discussion paper on generative AI devices, its first move toward regulating open-ended, non-deterministic clinical AI. The old framework assumed a locked model with a fixed output space; generative AI has neither. Validated will have to mean a narrow intended use, constrained and grounded outputs, a predetermined change-control plan, and continuous real-world monitoring, not a one-time test against a labeled set.

The FDA just signaled that the era of regulating AI as if it were a locked box is ending. STAT reported this month that the agency released a new discussion paper on generative AI devices, its first real move toward a framework for the open-ended, non-deterministic models now flooding into clinical software. If you build clinical AI, this is the shot across the bow you have been waiting for, whether you knew it or not. The word validated is about to mean something very different, and the teams that understand how will have a far easier next two years.

Why generative AI breaks the old rulebook

The existing framework for AI as a medical device was built around a very specific kind of model: locked, narrow, and deterministic. Think of a tool that takes a defined set of inputs and returns a risk score. You test it once against a labeled dataset, you freeze it, and if you want to change it you either file again or work within a predetermined change control plan. That approach is sensible, and it falls apart the moment the output is a paragraph instead of a number.

DimensionLocked ML device (old rules)Generative AI (new problem)
Output spaceFixed, such as a risk scoreOpen-ended text or plans
DeterminismSame input, same outputSame input, varying output
ValidationTest once against a labeled setMust cover an open output space
Failure modeMiscalibration, a wrong scoreFluent, confident fabrication
Change controlFreeze the modelModel and prompts evolve constantly

Every row in that table is a place where a one-time test stops proving anything. You cannot exhaustively test an open output space, a generative model can say something it never said in evaluation, and the scariest failure is not a slightly wrong number, it is a fluent, authoritative sentence that happens to be false. And this is not happening in isolation. The same season brought faster CMS coverage pathways for breakthrough devices and new scrutiny of how AI gets paid. The free-for-all phase is closing from several directions at once.

What validated has to mean now

If a single test can no longer carry the weight, validation has to become a system rather than an event. I do not need to see the final FDA guidance to know the shape it will demand, because it is the same shape good engineering already points to. Four moves do most of the work.

01

Narrow the intended use

The smaller and more specific the claim, the more testable it is.

02

Constrain the output

Structure it, ground it in sources, and refuse out-of-scope questions.

03

Predetermine change control

Specify up front how the model and prompts may change, and the bounds.

04

Monitor in the real world

Validation continues after release, not just before it.

Notice that narrowing the intended use is first, and it is the highest-leverage move by a distance. A generative tool that claims to do one specific thing for one specific population is something you can actually bound, test, and defend. A tool that claims to be a helpful clinical assistant for anything is, from a regulator's chair, an infinite surface of untested behavior. The broad claim feels like a bigger product. It is really a bigger liability.

What I would do now, before the guidance lands

A discussion paper is not a final rule, and that is exactly why now is the moment to move. You get to build the muscles the guidance will require while you still have slack, instead of retrofitting them under a deadline. Here is the order I would work in.

Now
Write the intended-use statement
Define the narrow claim you would be willing to defend to a regulator.
Next
Constrain and ground the output
Structure outputs, cite sources, and refuse out-of-scope questions.
Soon
Stand up post-market monitoring
Track performance and drift after release, with a change log.
When guidance lands
Adjust, do not scramble
You are refining a system you already built, not starting over.

There is a real competitive edge hiding in here, and it is not the edge people expect. The teams that treat regulation as an enemy will keep building broad, ungrounded assistants and then panic when the rules arrive. The teams that treat it as a design constraint will ship narrower, more defensible products that are, not coincidentally, also safer and easier to trust. In healthcare those are usually the products that win anyway.

For a decade, validated meant we tested the locked model once. For generative AI it has to mean we constrained it, and we never stopped watching. Build for the second definition now.

Naveen Kumar

I do not read the FDA moving on generative AI as a threat to the field. I read it as the field growing up. The discussion paper is an invitation to help define what safe looks like before the definition is imposed. Take it. Narrow your claims, constrain your outputs, write down how you will watch the thing after launch, and you will be ready for whatever the final guidance says, because you will have built for the only definition of validated that ever made sense for a system that can always surprise you.

Key takeaways
  • The FDA released a discussion paper on generative AI devices, its first real step toward regulating open-ended clinical AI.
  • The old rulebook assumed a locked, deterministic model. Generative AI has an open output space and varies run to run.
  • Validated has to be redefined: narrow intended use, constrained and grounded outputs, change control, and ongoing monitoring.
  • The failure mode also changes, from miscalibration to fluent, confident fabrication, which needs different tests.
  • Build for the new definition now. Write the intended-use statement and monitoring you would defend to a regulator.

Frequently asked

What did the FDA actually release?

A discussion paper on generative AI devices, reported in late August 2026. A discussion paper is not a final rule; it signals the agency's thinking and invites input before formal guidance. Treat it as the direction of travel, not the destination.

Why can't generative AI use the existing AI framework?

The existing framework grew up around locked models with a fixed output space and deterministic behavior. Generative AI produces open-ended output and can vary run to run, so a one-time test against a labeled set does not establish safety the same way.

What is a predetermined change control plan?

An FDA mechanism that lets an AI device be updated within pre-specified bounds without a new review each time. For generative AI, where prompts and models evolve constantly, defining those bounds up front becomes central to staying validated.

How does the failure mode differ?

A locked model tends to fail by miscalibration, a wrong score. A generative model fails by fabrication, a fluent and confident wrong answer. That is harder to catch and needs grounding, output constraints, and monitoring rather than a single accuracy check.

What should product teams do before formal guidance?

Narrow the intended use, constrain and ground the outputs, write the model card you would defend, and stand up post-market monitoring. When the guidance lands you will be adjusting a system you already built, not starting from zero.

Sources

Naveen Kumar

Naveen Kumar

Healthcare engineering and product executive in Pittsburgh. 15+ years building AI-first patient access, a decade at Treatspace.

Read next