AI BioDesign launches with $94.6 million to build proteins beyond nature

AI BioDesign launches with $94.6 million to build proteins beyond nature

N
News Editor
2026-09-28 02:02:08
AI BioDesign, a new institute launched in Seattle on Sept. 3, is taking a different path from AlphaFold-style biology. Backed by $94.6 million from the Paul Allen estate and led by 2024 Nobel chemistry laureate David Baker and genomics researcher Jay Shendure, the group is focused on designing molecules that nature never evolved rather than predicting structures that already exist. Its core idea is not to spend the money on a larger general-purpose model, but on experimental throughput. The institute runs a loop in which AI proposes designs, the lab builds and measures them, and the results are fed back into the model for the next round. In one recent experiment, researchers tested 6 million designed DNA sequences in a single tube, a scale the article says would have taken years with traditional methods. The group argues that the bottleneck in protein design is no longer generating candidates but validating them fast enough. It also says its models, datasets, methods, reagents, and benchmarks will be open. At the same time, the effort raises biosafety questions around synthetic DNA screening, which Baker says should be addressed with sequence registration and purchaser records.

AlphaFold has predicted the structures of more than 200 million proteins found in nature. David Baker is chasing something else entirely: not guessing what evolution already built, but making molecules evolution never made.

AI BioDesign launches with $94.6 million to build proteins beyond nature 2

One morning in August, research assistant Jack Boylan opened the results from an overnight run. Beside a DNA sequencer sat test tubes packed with artificially designed DNA sequences that do not exist in nature. The machine had one job: read them all in one sweep and separate the useful designs from the failures.

In a recent experiment, the team loaded 6 million of those sequences into a single tube and measured them together. With a standard setup, the article says, handling that batch by itself would have taken years.

The group doing this is AI BioDesign, which stepped into public view in Seattle on Sept. 3, a few minutes from the Allen Institute headquarters. It has $94.6 million in backing from the Paul Allen estate system and is led by Baker, a 2024 Nobel Prize in Chemistry laureate, together with genomics scientist Jay Shendure.

At a moment when plenty of labs are pushing toward bigger models, this group is not about to spend that nearly $100 million on training a larger one. Its bet is the test tube. Simple as that.

Parallelism inside the tube

AI BioDesign is building a rapid loop: AI suggests designs, the lab makes and measures them, the results go back into the model, and the model decides what to make next.

On paper, that sounds familiar. It is still a design-build-measure-learn cycle. But the scale changes everything. There is no factory floor stuffed with robots. Executive director Jesse Gray put it bluntly: "The scale does not come from robots. It comes from doing the parallelism inside the test tube."

Shendure describes it as taking the cellular production line shaped by evolution and repurposing it to make and test millions of molecules. The lab drops millions of designed sequences into one tube, lets cells produce them, runs them under identical conditions, and then sends the output to a sequencer that reads, in a single pass, which ones worked and which ones failed.

AI BioDesign launches with $94.6 million to build proteins beyond nature 3

That gives the model a huge training set of examples about the rules of design, rather than forcing it to learn from the relatively small collection of molecules nature happened to leave behind.

The article points to an earlier Baker lab design as an example: a binding protein in which the purple and pink regions were designed by AI, while the white region was amylin wrapped by that designed protein.

Experiments are judged by what the model learns

The production line is scored differently, too, compared with a standard biology lab. Teams of five or six are assigned to a specific design problem, and a machine learning group works right alongside them.

Each experimental round is not judged only by how many useful molecules it spits out. The institute also asks how much the model learned from the results.

In a normal lab, researchers begin with a hypothesis and then build experiments to test it. AI BioDesign flips that around. First, experiments are framed as practice problems for the model. Then the design-build-measure-learn cycle keeps spinning.

After every round, the machine learning team picks out the sequences the model is least certain about and the ones most likely to add fresh information, then uses that set to decide the next measurements. Those data go back into the model, which updates and generates the next batch of designs.

The aim is to squeeze as much information as possible out of each experiment, not to rush a finished product out the door. In the article’s wording, the bench used to follow people; now it is starting to follow AI.

Baker was designing non-natural proteins long before this launch

This is not some first-ever story about AI making a protein nature never produced. Baker did that 23 years ago.

AI BioDesign launches with $94.6 million to build proteins beyond nature 4

In 2003, his team used Rosetta software to design a protein called Top7. It had 93 amino acids, and its folding pattern had no precedent among the proteins known at the time.

In the 2024 Nobel Prize in Chemistry, one half went to Baker for computational protein design. The other half went to Demis Hassabis and John Jumper of the AlphaFold team for protein structure prediction. On Oct. 9, 2024, David Baker received the call from the Nobel committee.

If AlphaFold solved the reading problem — predicting what shape an amino acid sequence will fold into — Baker’s side is the writing problem: deciding what shape should be built to produce a chosen function.

In recent years, the reading side moved to the center of the field. AlphaFold2 and other large models from that same era were trained on data humanity had already gathered. The writing side does not get that luxury. Molecules that do not exist in nature are not waiting in a database. They have to be made and measured from scratch.

That is why the field has remained stuck under a practical ceiling. AI can produce hundreds of thousands of candidate designs overnight. A traditional lab might make only a few hundred in a year. That mismatch — generation on one side, validation on the other — has been the real bottleneck.

Baker said the real breakthrough here is that AI speed is beginning to line up with experimental capacity in synthetic biology. One side can generate millions of candidates in a day. The other can test millions in a tube. At last, those two gears are meshing.

Not a general model bet, but an infrastructure bet

While others are chasing a general model that can answer everything, or maybe even simulate an entire virtual cell, Baker and AI BioDesign are going another way. They plan to build smaller models aimed at specific problems, then generate large volumes of new data for each one.

AI BioDesign launches with $94.6 million to build proteins beyond nature 5

The logic is straightforward. The samples left by natural evolution are only a small slice of what billions of years happened to explore. Try to infer broad design rules from that slice, and a big gap remains. So you fill the gap yourself by generating the missing data.

And that depends less on piling up GPUs than on tubes, sequencers, and lab benches. The article frames it this way: the last phase was a race for compute infrastructure. The next one may be a race for experimental infrastructure.

Even stranger, the payoff is supposed to be open. Models, datasets, experimental methods, reagents, and benchmarks will all be released publicly.

Asked what happens if others use those data to build profitable drugs, president Rui Costa answered lightly: "If many companies use these data to make the world better, that is our good fortune."

Biosafety concerns rise with design power

The article also points to the risks. Making molecules that do not exist in nature at scale raises the possibility of unintended ecological effects and malicious misuse.

One step in that chain cannot be avoided: synthetic DNA. No matter how many designs a model proposes, someone still has to turn them into physical material. Right now, the gatekeeper is DNA screening. When a customer orders a synthetic DNA sequence, the system compares it against known toxin and pathogen genes and blocks the order if it finds a match.

A study published in Science in October 2025, the article says, showed that a Microsoft team used public protein design tools to rewrite 72 proteins of concern into about 72,000 synthetic homolog sequences. Those kept broad structural and functional properties while changing the sequence substantially, and four screening tools missed many of them.

The Microsoft team later worked with four synthesis companies to patch the problem, and detection rates improved sharply.

AI BioDesign launches with $94.6 million to build proteins beyond nature 6

Still, a patch fixes only the weakness that has already been found. Design capabilities keep moving ahead. The gatekeepers are still chasing from behind.

Baker’s proposal is that all artificially synthesized DNA should be registered and archived, with both the sequence and the identity of the purchaser recorded. He argues that this would create a practical barrier against misuse.

A timeline of 18 to 24 months

The bet is already live. Costa set a target of showing real progress within 18 to 24 months and opening the platform to researchers worldwide within five years.

Baker’s stated ambitions include therapies for new diseases that take weeks rather than decades, enzymes that can consume plastic in the ocean, and biological computers that use far less energy than silicon chips.

For now, those are still directions, not finished outcomes. Baker also acknowledged that the first results may be molecules that work even when researchers cannot yet explain why they work.

Darwin wrote of "endless forms most beautiful" to describe nature’s variety. Baker’s view, as the article presents it, is that those beautiful forms are only a small fraction of what billions of years happened to test. Most of the drawers are still shut. With AI, he is trying to yank them open, 6 million compartments at a time.

AlphaFold has already gone a long way down the path of reading nature by predicting the structures of more than 200 million proteins. Baker’s wager is that the next bottleneck is not the model itself, but data — and the experimental pipeline needed to keep making more of it.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.