AI Health Forecasting Is a Data Problem First
The models that forecast your health are real. The decades of clean data they need are the actual problem.

Photo by Marek Piwnicki on Pexels
My take: the exciting part of AI health forecasting is the model, but the hard part is the data it demands. Predicting a person's health arc needs decades of longitudinal, well-structured records, and most health systems do not have that. The forecast is only as good as the history behind it.
Every few weeks a new model claims it can predict your health years out, and the demo is stunning. My take, after a decade building on health data, is that the model is the easy half of the story. The hard half is the data it quietly assumes you already have. A forecast of a person's health arc needs a long, clean, linked history, and most health systems simply do not have one. The forecast is only ever as good as the record behind it.
The model is real, and it is impressive
The latest forecasting models are not hype, and I want to be honest about that before I push back. I read Eric Topol's Ground Truths interview with Sarah Urbut about a model called ALADYNOULLI. It does something new. Instead of scoring one disease, it forecasts trajectories across 348 conditions at once, learning what the authors call latent disease signatures and folding in genetic risk. The method is written up in Nature as a Bayesian framework for longitudinal records and genetic discovery. This is a real step, not a press release.
The gap is not small. On a one-year coronary prediction the model reached an AUC of 0.89, against about 0.68 for a standard risk score. That is the difference between a useful warning and a coin flip.
Now read the fine print on the data
The number that should stop you is not the accuracy. It is the data the model ran on. It learned from about 683,000 people across three cohorts, with follow-up stretching up to 52 years, covering 348 diseases at once. Sit with that. Fifty-two years of linked, coded, usable history on the same people. That is the real prerequisite, and it is where almost every health system I have worked with falls short. We have plenty of data. We do not have plenty of long, clean, linked data on one person over decades.
What forecasting actually demands of the rest of us
If you want models like this to work on your population, the hard work is not the model. It is years of unglamorous data discipline. Keeping it general, four things have to be true. You need to follow a person across time without losing them when they change plans or sites. You need diagnoses coded consistently across that whole span, not redefined every few years. You need to link records across settings into one timeline. And you need to treat missing data as information, not as zero. None of that is AI work. It is data work, and it has to happen first.
Where I land
For most organizations, the honest first move toward AI forecasting is not a model. It is a decade of disciplined data, or a partnership with someone who already has it. I am not against the ambition. I am for sequencing it right. If your longitudinal data is thin, a forecasting model will hand you confident predictions built on a shaky history, which is worse than no prediction at all, because people will trust it. So my advice is plain. Fund the record before you fund the forecast. Build the long, linked, consistently coded history first. The model will still be there when your data is ready, and it will be better than the one you would have shipped too early.
- AI health forecasting is a data problem before it is a model problem.
- New models forecast across hundreds of diseases at once and beat standard risk scores by a wide margin.
- That accuracy was earned on about 683,000 patients with up to 52 years of linked follow-up, a bar most systems cannot meet.
- Forecasting your population needs longitudinal linkage, consistent coding over time, cross-setting record linkage, and honest handling of missing data.
- A confident forecast on thin history is worse than no forecast, because people trust it.
- Fund the record before you fund the forecast.
Frequently asked
What is AI health-trajectory prediction?
It is using models to forecast a person's likely future health across many conditions over time, rather than scoring a single disease risk at one moment.
Why is data the harder part?
Because these models learn from long, linked, consistently coded histories. Most health systems have large data but not decades of clean, connected records on the same people.
How much data did the model Topol profiled use?
It trained on about 683,000 patients across three cohorts, with follow-up up to 52 years, modeling 348 diseases at once.
How accurate is it?
On a one-year coronary prediction it reached an AUC of 0.89, compared with roughly 0.68 for a standard risk score.
What should an organization do first?
Build longitudinal data discipline, consistent coding, record linkage, and honest handling of missingness, before funding a forecasting model. A confident prediction on thin history is worse than none.