According to Dongcha Beating, Fudan University Vice President Jiang Yugang, AgiBot partner Yao Maoqing, Tashi Zhihang CEO Chen Yilun and Liangyuan Xinchuang CEO Jiang Xu took part in a roundtable discussion at the 2026 World Artificial Intelligence Conference, focusing on world models. Their shared view was that general-purpose embodied intelligence is still some distance away, and that the field will need to break through in specialized scenarios before it can move broader. They also said the center of competition is likely to move away from model architecture alone and toward access to high-quality data and the ability to validate systems in closed-loop environments.
World models were framed as tools for understanding the physical world
The speakers said the key job of a world model is not image rendering. In their view, it is about understanding how the physical world operates and predicting the next state or action. That requires native capabilities in multimodal fusion, physical-law modeling, causal reasoning and long-horizon prediction, rather than visual generation alone.
Data was described as the biggest bottleneck
Chen Yilun said current video datasets lack critical modalities such as force and touch, which limits their value for embodied AI training. He said ideal training data would need to meet three conditions: full modality coverage, high-frequency interaction and roots in real-world environments. Given the complexity of embodied tasks, he added, the field may need tens of millions of hours of real interaction data.
Yao Maoqing compared the challenge with large language model training, saying that those systems have been supported by tens of billions of hours of speech training data. Based on that comparison, he estimated that learning common-sense physical prediction may require more than 100 million hours of real-world data.
Mainstream architectures were said to mix state and action prediction
On model design, Jiang Xu said current mainstream architectures tend to handle state prediction and action prediction together. In his view, that creates a conflict between generation and understanding, making it difficult to optimize both at the same time.
Manufacturing was seen as the clearest scaling path over the next three years
When the discussion turned to deployment, the speakers pointed to manufacturing as the most certain large-scale application scenario over the next three years.
Yao Maoqing said AgiBot has already deployed robot swarms on production lines, reaching 60,000 operations in six days with a 99.99% success rate.
Chen Yilun said he is also betting on manufacturing. He cited three reasons: higher data density, clearer task completion standards and large volumes of human demonstration data. Chen added that Tashi Zhihang is working with automakers on the deployment of industrial embodied robot clusters at the thousand-unit level. He also said China has the world’s most concentrated manufacturing base, making it an ideal testing ground for physical AI.
Homes and offices may see earlier capability jumps
Jiang Xu said embodied intelligence should be seen as an extension of multimodal foundation models. He noted that the internet already contains 10 billion hours of video data that is suitable for pretraining, and said major capability jumps may first appear in everyday settings such as homes and offices. Even so, he argued that commercialization requires high fault tolerance, and that finding the right scenario for a large model is no easier than training the model itself.
Panel consensus favored a narrow-to-broad path
The common conclusion from the discussion was that general-purpose embodied intelligence remains far off. Progress will have to come through specialized scenarios first, while the next phase of competition is set to focus on who can secure better data and who can complete closed-loop scenario validation.

