Chan Zuckerberg Biohub said on Oct. 7 that Google DeepMind, Isomorphic Labs, Meta, the U.S. Department of Energy and the National Institutes of Health have joined its Virtual Biology Initiative, a project focused on building open data for future “virtual cell” AI models. Biohub put the total value tied to the effort at $1.8 billion.
The group made clear that this is a data infrastructure program, not a model launch. AI systems that can accurately predict how cells behave still do not exist.
Data is the main constraint
A virtual cell, in simple terms, would be an AI model that predicts how a cell responds to drugs, mutations or disease. The idea is to let scientists run more experiments on computers first, then reserve expensive lab work for the most promising questions.
The problem is data. Progress in other AI fields has been driven by stronger computing power and large existing datasets. Biology is different because the data has to be measured from the real world, one observation at a time.
Biohub science lead Alex Rives told Reuters that current cell datasets are on the order of hundreds of millions of cells. Accurate predictive models will need billions, and eventually trillions.
“We need to capture the language of biology, capture the language of cells, and that language does not exist today,” he said.
In an April release, Biohub compared the effort to the Protein Data Bank and the Human Genome Project. The first became one of science’s most important resources through contributions from researchers around the world. The second required top labs globally to align around a shared goal.
Biohub background
Biohub was founded in 2016. The article says that in November 2025, the Chan Zuckerberg Initiative moved its science operations under Biohub and folded in AI biology lab EvolutionaryScale for an undisclosed amount, with co-founder Rives becoming science lead.
Zuckerberg said at the time that the original goal was to cure, prevent or manage all disease by the end of the century, and that advances in AI meant it “may happen much faster.”
How the $1.8 billion breaks down
The headline figure looks less dramatic once broken into parts.
- Biohub itself accounts for $500 million. Of that, $400 million is for internal measurement technology and $100 million supports external research. That commitment was announced on April 29 and is not new.
- The U.S. Department of Energy is contributing more than $500 million over five years through the DOE-led Genesis Mission for lab measurement, modeling and computing.
- NIH is not adding new money under this announcement. In Biohub’s wording, NIH is coordinating datasets, databases and knowledge bases built from more than $500 million in past federal investment, and Biohub will help standardize them for AI training. NIH’s own announcement did not mention any dollar amount and instead described integrating existing biomedical datasets and national data infrastructure.
- Google DeepMind, Isomorphic Labs and Meta are contributing a combined $300 million. Biohub did not disclose how much each company is putting in.
The article identifies Isomorphic Labs as an Alphabet-owned AI drug discovery company that was spun out of DeepMind.
Open data comes with a one-year exclusivity window for commercial partners
Rives told Axios that commercial partners will get a one-year exclusivity period for data generated from their funding before it is released publicly.
“We have to give commercial participants some incentive to join, and the data embargo period is what provides that. But the embargo period is one year,” he said.
He also told Reuters that government-funded data will not face that restriction. The next step, he said, is to approach pharmaceutical companies and philanthropic groups.
The timeline is an expectation, not a delivery commitment
On timing, Rives told Reuters that work like this would normally take decades, but the partners want to compress it into five years. The first datasets could be ready in about a year, and he expects accurate predictive models within five years. The article notes that these are his expectations, not signed delivery deadlines.
Rives had already raised a central question with Axios in April: whether cell biology will show the same kind of scaling law seen in other AI fields, where feeding models more data leads to predictable gains. In his latest comments to Axios, he said that within a year of getting the first large-scale datasets, researchers should be able to train models, measure their capabilities and identify what additional biological data would improve them.
Pointing to AI progress in protein biology, he said: “It works in every domain, and it works in biology too.”
Others are also expanding AI biology efforts
Priscilla Chan told Reuters: “We’ve always thought of this as a shared asset for the whole community, not something that belongs to any one group, so that it can keep compounding over time.”
Reuters also noted that Anthropic is increasing its investment in biology and building its own wet lab, while the OpenAI Foundation has launched a biomedical dataset grant program worth more than $125 million.

