In a striking departure from the prevailing trend in artificial intelligence, former OpenAI researcher Andrew Ho is predicting a major shift in how AI models are developed. Ho, who is now launching his own company, argues that the current approach of simply scaling up language models is no longer sufficient to drive meaningful progress. Instead, he believes that the future lies in specialized training data, a move that could see more than $100 billion invested in targeted data collection efforts.
The Decline of Generalist AI
Ho’s concerns are echoed by fellow researcher Adam Hunt from Cambridge. Together, they observe a troubling trend in the AI landscape: models that are becoming increasingly specialized, excelling in narrow domains like coding and mathematics, but losing ground in broader, more general tasks. This specialization, they argue, is a sign that the era of relying purely on scale to improve performance may be coming to an end.
Why Scaling Isn’t Enough
While large language models have historically grown in size and complexity, the gains in performance have not always translated into better real-world capabilities. Ho and Hunt believe that the solution lies not in bigger models, but in better, more curated data. By focusing on high-quality, domain-specific datasets, AI labs can train models that are not only more accurate but also more adaptable to specific tasks.
The Road Ahead
Ho’s new venture is set to capitalize on this insight, aiming to revolutionize how training data is collected and used. As AI labs grapple with the limitations of scaling, the shift toward specialized datasets could redefine the industry’s priorities and investment strategies. With $100 billion potentially at stake, the move toward targeted data collection may well become the next big wave in AI development.



