Data in Community and Social Care: The Hardest Data to Get Right
The Healthcare Data sub-pillar about the conditions of daily life that shape health most. Why this data is the hardest and most human, and how to handle it with care.
Data in community and social care is the frontier of healthcare data. It covers the conditions of daily life that shape health more than the clinic does, and it is the thinnest, most fragmented, and most sensitive data we have. In my experience the hardest part is not collecting it. It is remembering that the people missing from the data are usually the ones who need help most.
Data in community and social care is the hardest problem in my healthcare data pillar, and I put it last on purpose, because it is the one you should approach only after you respect how hard it is. It covers the conditions of daily life, housing, food, income, isolation, transport, the things the World Health Organization calls the social determinants of health. These shape health more than the clinic does. And the data about them is thin, scattered across systems that were never meant to talk, and full of people who are barely recorded at all.
So here is where I land before writing a line of code. This data can do real good, connecting people to support they need. It can also do real harm, if you flatten a person into a risk score or treat a gap in the record as a zero. I treat community and social care data as the most human data we hold, and I handle it accordingly: carefully, at the level of principle, and with the people in it firmly in mind.
What the social determinants are
Most of what determines a person's health happens outside the healthcare system, in the conditions the WHO calls the social determinants of health. The WHO describes them as the conditions in which people are born, grow, live, work, and age, along with their access to power, money, and resources. Research consistently finds these can outweigh both clinical care and genetics in shaping outcomes. They are commonly grouped into a handful of domains, and each one is a different data problem, because each lives in a different system, if it is recorded at all.
The missing data problem
The most important fact about community and social care data is that the gaps are not random: the people with the least data are often the ones with the greatest need. Think about who does not show up cleanly in health data. Someone without stable housing, without steady insurance, moving between clinics and none, falling through the seams of systems that do not connect. The result is a data shadow. If you train a model, or write a report, on the people you can see, you systematically underserve the people you cannot. This is the trap I watch for most, because it hides inside data that looks complete.
Close the loop, do not just collect
Collecting social data you do nothing with is worse than not collecting it, because you asked, and then failed the person who answered. The point of this data is action, not a richer dashboard. That means a full loop: screen for a need, record it in a shared, standardized way, connect the person to a real service, and confirm the loop closed. The Gravity Project exists precisely to standardize how social needs are coded and exchanged so that loop can work across clinical and social systems. Standardization matters here, but only in service of the referral. If you screen someone for food insecurity and there is no service on the other end, you have collected a sensitive fact and delivered nothing. I would rather not ask.
Principles I hold to
Because this data can help and harm in equal measure, I hold to a few hard principles before I touch it.
Treat every gap as a signal about access and reach, not as a zero to impute away. The absence tells you something, usually about who your systems are failing to see. Fill it with a default value and you erase the very people you should be looking for.
A social need shared in confidence is not data to spread across systems freely. I collect the minimum that lets me help, I am clear about why, and I guard it more carefully than clinical data, not less. The bar for using it should be higher, because the harm from misusing it lands on the people least able to absorb it.
Health systems and social services run on separate data worlds, and the person lives in both. Without connecting them, you only ever see half of someone. And a last caution I keep close: a model trained only on who you can see will quietly fail on who you cannot, so I test for that bias on purpose rather than hoping it is not there.
Data in community and social care is where healthcare data stops being a technical problem and becomes a human one. It rounds out my healthcare data pillar because you cannot understand a person's health from the clinic alone. But it asks for more care than any other data we hold. Handle it with consent, with humility about what is missing, and with a service on the other end of every question. Do that, and the data helps. Skip it, and you have built a very precise way to overlook the people who need you most.
- Community and social care data covers the conditions of daily life that shape health more than the clinic does.
- The WHO calls these the social determinants of health, and they can outweigh clinical care and genetics.
- The gaps are not random: the people with the least data often have the greatest need.
- Treat missing data as a signal, not a zero, and guard this data with extra consent and care.
- Close the loop: screen, code with a shared standard, refer to a real service, and confirm it resolved.
- A model trained only on who you can see will fail on who you cannot, so test for that bias.
Articles in this topic
More articles coming soon.
← Back to Healthcare Data: The Foundation Under Every AI and Product Decision
Frequently asked
What is data in community and social care?
Data about health needs and care outside the hospital, including social determinants like housing, food, income, and isolation, often held across separate health and social systems.
What are the social determinants of health?
Per the WHO, the conditions in which people are born, grow, live, work, and age, plus access to power, money, and resources. They can influence outcomes more than clinical care or genetics.
Why is this data so hard to work with?
It is thin, fragmented across systems that do not connect, sensitive, and systematically missing for the people with the greatest need.
What is the Gravity Project?
A national public collaborative that develops consensus data standards for how social determinants of health are coded, documented, and exchanged across clinical and social systems.
What does missing is informative mean?
That a gap in social data usually reflects who your systems fail to reach, so you should treat it as a signal rather than filling it with a zero.
How do you use this data responsibly?
Collect the minimum needed to help, get consent, guard it carefully, connect every screening to a real service, and test models for bias against the people underrepresented in the data.