OpenAI Foundation Launches Public Data for Health, Funding Biotech Data Troves for AI Drug Discovery
核心洞察
The OpenAI Foundation (搜索) launched Public Data for Health, an effort to pay for high-quality scientific datasets it says AI needs to make medical breakthroughs.
A $500,000 grant to 1Day Sooner (搜索) will fund bids at bankruptcy proceedings to acquire common technical documents from failed biotech companies.
The foundation also pledged $40 million for novel cancer (搜索) vaccine data collection at the University of North Carolina, Chapel Hill, and support for OpenAdmet (搜索).
The OpenAI Foundation (搜索), the nonprofit parent of OpenAI (搜索), announced it will fund an effort to acquire the regulatory filings, manufacturing strategies, and safety data of bankrupt biotechnology companies, part of a new initiative called Public Data for Health aimed at paying to create "high-quality scientific datasets" for artificial intelligence in medicine.
The idea originated with a proposal to bid at bankruptcy proceedings for detailed regulatory filings, manufacturing strategies, and safety data — information typically treated as trade secrets. The proposal's author, described as a writer for Works In Progress and a nonresident fellow at the Institute for Progress, called these documents "biotech's lost archive" and said they could be used to train AIs that would act as powerful copilots in the often opaque drug approval process.
The foundation awarded $500,000 to pursue the idea, and it will be carried out by 1Day Sooner (搜索), an advocacy group representing clinical trial volunteers that the proposal's author advises. The drug company files being sought are known as common technical documents, which typically contain the back-and-forth between companies and regulators, detailed scientific and medical measurements, and essentially everything that is known about a drug.
Josh Morrison, president and cofounder of 1Day Sooner (搜索), said the grant will help the group prove it can obtain the data troves of bankrupt companies. He believes nonexclusive copies of company datasets could be acquired for only "a few tens of thousands of dollars" each. Morrison noted that two other attempts to obtain drug company files this year proved unsuccessful after 1Day Sooner's bids were not accepted.
The proposal's author argues that a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market. "People say 'We will invent AI, and AI will cure cancer (搜索),' but that's very removed from the messy reality and the regulatory process," she said. "About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations."
Broader Data Grants
In its initial round of data grants, the OpenAI Foundation (搜索) also announced it would give $40 million to a program to collect data about novel cancer vaccines (搜索) at the University of North Carolina, Chapel Hill, and support OpenAdmet (搜索), a group that runs competitions in which researchers try to predict drug effects.
The foundation framed the effort around a data bottleneck. "Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology," said Morgan Levine, a former vice president for computation at Altos Labs, a longevity company. The foundation said in a statement: "We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data."
Jacob Trefethen, an executive at the foundation, said it essentially operates separately from OpenAI (搜索) but shares an official mission of ensuring that artificial intelligence "benefits all of humanity." He said the foundation hopes to give away $1 billion by the end of the year. Based in San Francisco, the foundation is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition (搜索), an organization that helps patients get access to drugs for hepatitis C (搜索).
OpenAI (搜索) started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion. Because the foundation holds a 26% equity stake in OpenAI, it is on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.
The charitable push comes amid broader debate about AI risk. The report notes that apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon, with some AI company insiders saying the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to "slow the pace at which we improve the capabilities of AI models" so that risk prevention can catch up.
Bankruptcies could become what some are calling a "new land grab" for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails, which led to objections from flight attendants and others who worried that private or proprietary data could be exposed.
A Parallel Bet on Paying Data Generators
A separate startup is testing a related premise: that the scientists and companies generating high-quality biological data should share in the value created from it. Harell Data (搜索), based in Bellevue, WA, announced it has raised $15 million from Fuse, Cercano Management, and others to build what it describes as a more sustainable and mutually beneficial cloud computing service for data generators and AI model builders.
The company was founded by Harlan Robins, scientific founder of Seattle-based Adaptive Biotechnologies, a public company worth $3.9 billion. "Data itself, especially for training ML models, isn't getting appropriately valued," Robins said. "Everyone gives it lip service, how the important the data is. But all models generate huge value, whether it's ChatGPT, Claude, or others, from training on public data."
Harell has lined up secure cloud computing services through CoreWeave Cloud (搜索) with Nvidia (搜索) graphics processing unit technology, according to Robins. Those computers are initially being populated with two well-curated, proprietary datasets: one from Seattle-based A-Alpha Bio (搜索) covering protein-protein binding affinities (搜索), and one from Adaptive covering proprietary T cell receptor (搜索) binding data.
Companies that provide data to the scientific community for drug discovery will receive a cut of the revenue Harell receives from AI model builders, Robins said. AI model builders, he said, will benefit by getting access to well-curated datasets relevant to drug discovery, and the price to customers will be lower than what they currently pay to big cloud service providers such as Amazon Web Services, Microsoft's Azure, and Google Cloud.
Robins said he tried to circulate the idea among the big cloud computing companies, offering to drive hundreds of millions in revenue to them if they would offer a cut of revenue to data generators. Those talks went nowhere. "I understand why," Robins said. "They are selling every bit of compute they are getting already. They are supply limited, not limited on customers."
Robins is not disclosing the percentage cut of revenue Harell offers data generators, but said they will be paid upfront and the company will not try to reach through for downstream royalties on discoveries years after the fact. Harell has hired a team of 11 people, none brought over from Adaptive, where Robins remains a consultant.
Robins pointed to the Protein Data Bank as the unsung hero that made AlphaFold possible, noting that for decades structural biologists, biochemists, and others who put carefully curated protein structures in the public domain laid the conditions for AI to predict protein structures — and never got a nickel. He said much industry data lives in silos and is not accessible to AI model builders, and that the only way to find out whether a different business model can accelerate discovery is to try it.
