Andrew Ho, a researcher who worked on OpenAI’s biomedical AI evaluations including GeneBench and GeneBench-Pro, says he has left OpenAI and is starting a new company focused on high-quality reinforcement learning, or RL, training datasets. His view is that the next bottleneck for large language models is no longer just adding GPUs or scaling parameter counts, but getting data that can materially improve reasoning.
Ho says LLM generalization is still weak
In his departure post, Ho argues that the public still misunderstands what current large models can do. He says LLM generalization remains limited. Even after billions, and in some cases hundreds of billions of dollars have gone into AI development, models still display what he describes as “spiky capabilities” across many fields.
He points to software work as an example. Current models may solve Codeforces problems and even outperform him at porting C++ code to Rust, he says, but he still has to manually revise AI-generated pull requests in real-world work. For Ho, that gap shows that models are still far from truly understanding software development. In his telling, today’s systems are better described as excelling on a narrow set of benchmarks than as having broad, reliable generalization.
The shortage, in his view, is RL data rather than models
Ho argues that the bigger market problem is not the model itself but the shortage of high-quality RL data. He says most existing data vendors have not actually participated in large-model training, and as a result they do not know what kinds of data can improve model capability in a meaningful way.
He also says many economically valuable tasks are deeply context-dependent. Business decisions, medical judgment, and scientific analysis are hard to turn into RL environments with clean scoring in the way math problems can be. Even if a human follows a so-called “Golden Path,” it can still be difficult to judge whether other decision paths are good or bad. He describes that as one of the core challenges in RL training today.
He predicts more than $100 billion could go into data purchases
Ho makes a broader prediction as well. Over the next few years, he says, leading AI labs could spend more than $100 billion buying high-quality training data instead of relying only on ever-larger model scale.
His reasoning is that companies such as OpenAI, Anthropic, and Google are coming under growing pressure to make money, while scaling laws on their own are no longer enough to keep lifting model performance. What matters next, he says, is precise and highly targeted data acquisition. Ho adds that he directly handled the full process at OpenAI, from data procurement to model training, and says that experience gives him a clear view of what large AI companies actually need from data suppliers. He wants his new company to become one of those suppliers.
The first focus is biology and statistical reasoning
Ho says the company’s first products will target two areas. The first is datasets built for long-horizon scientific reasoning. He says GeneBench-Pro, which he recently released with the OpenAI team, was designed to test whether AI can actually carry out research analysis that involves repeated judgment, exploration, and revision, rather than just answering multiple-choice questions.
He says GPT-5.6 Sol currently reaches only about a 30% top pass rate on GeneBench-Pro. In his view, that leaves a long distance before models can reliably assist scientific research. His goal is to use high-quality RL data to raise model reliability above 90%.
The second direction is data that reflects routine research work more closely. Ho gives examples such as a researcher taking a photo of a cell culture plate, a Western Blot, or another experimental image and asking AI how to interpret it. He says current systems, including multimodal models, still perform poorly on these real scientific tasks. If AI is supposed to help with scientific discovery, he argues, those basic capabilities have to be built first.
Later expansion could reach healthcare, materials, and white-collar work
Beyond biology, Ho says the company plans to expand gradually into RL data for chemistry, materials science, healthcare, and even general white-collar work. He also addressed AI labs including OpenAI and Anthropic directly, saying the company will offer data services under industry-standard pricing and partnership structures and inviting interested AI companies to contact him.

