NIH Unveils World's Largest Genomics-and-Health Database, Linking 535,000 Whole Genomes to Medical Records
核心洞察
The NIH's All of Us programme released data from over 747,000 participants, including 535,000 whole genome sequences linked to nearly 482,000 electronic health records, creating the world's largest integrated genomics-health database.
More than 86% of participants come from groups historically underrepresented in medical research, including racial and ethnic minorities, enabling discoveries that narrower databases would miss.
The resource has already fueled over 1,400 peer-reviewed papers and contributed to genetic risk prediction tools across cardiovascular conditions (搜索), prostate cancer (搜索), and Alzheimer's disease (搜索).
On June 30, 2026, the National Institutes of Health released data from more than 747,000 participants through its All of Us research programme (搜索), creating what it describes as the world's largest integrated store of genomes and electronic health records. The release links 535,000 whole genome sequences to nearly 482,000 medical records — a combination of genomic depth and clinical breadth that the NIH calls "unmatched by any research programme in the world."
The trove now holds more than 1.3 billion genetic variants and bundles genomes with doctors' notes, diagnoses, test results, health surveys, wearable data from devices like Fitbits, and even local air quality measurements. The eventual goal is to enroll one million volunteers.
A database built for those left out
What distinguishes All of Us from existing resources is its participant composition. More than 86% of participants come from groups long overlooked in medical research, including racial and ethnic minorities, older adults, women, people with disabilities, and rural residents spanning all 50 states.
This stands in contrast to the UK Biobank, the field's leading repository with records for approximately 500,000 people, almost all of whom are of white European ancestry. Findings from such homogeneous datasets often fail to generalize to other populations. All of Us has already helped uncover gene variants that reduce the risk of kidney disease (搜索) in people of African ancestry — a result a narrower database would not have yielded.
"One of the most exciting components is its sheer diversity," said Alicia Martin, a statistical geneticist at the Broad Institute who uses the data to build risk-prediction tools. She noted the resource offers a way to understand not just who is at risk of disease, but who will respond to which treatment.
Fueling AI and the multiomics era
The database arrives as modern drug discovery and diagnostics increasingly depend on large, clean datasets. A genomic-plus-clinical trove at this scale provides precisely the raw material that new AI research tools and science-specific models require. The release also inaugurates what the programme calls its multiomics era, adding protein and RNA data for thousands of participants.
The payoff is already tangible. All of Us data has fed more than 1,400 peer-reviewed papers from approximately 23,000 researchers. It helped build a genetic test that predicts inherited risk across eight cardiovascular conditions (搜索) and supported a low-cost prostate cancer (搜索) model now in a trial of 5,000 veterans. Early work on catching disease before it starts, including Alzheimer's, has also drawn on the resource. The programme has returned more than 733,000 personalised DNA results to participants. Access is free, ensuring a researcher at a rural university receives the same data as one at a top institute.
A national treasure under threat
The milestone is shadowed by financial uncertainty. The programme's budget has been cut by 72% since 2023, and one of its main funding streams, the 21st Century Cures Act, is set to expire at the end of the current fiscal year. More than 50 medical organisations have written to Congress warning that much of what has been built could be lost.
NIH Director Jay Bhattacharya called the database "a national treasure" and a foundational platform for investigators at every career stage. Both characterizations hold true simultaneously: the resource is more valuable than ever, and its future is less certain than ever.
The privacy foundation
A repository of half a million genomes linked to medical records represents one of the most sensitive datasets ever constructed. All of Us shares it only with registered researchers through a controlled cloud workbench, relying on the trust of the people who volunteered their data. As AI makes it easier to draw conclusions from health data, guarding that trust will only grow more challenging.
For now, scientists have never had a health map this large or this representative of the actual population. Whether the country continues to fund it will determine how much of that promise becomes reality.
