Keyu Tian’s world-model startup reaches $200 million valuation before launching a product

Keyu Tian’s world-model startup reaches $200 million valuation before launching a product

N
News Editor
2026-10-08 08:53:27
Researcher Keyu Tian, who drew attention after a dispute during his internship at ByteDance, is back in the spotlight as the founder of a new AI lab focused on world models. Bloomberg reported on Oct. 7 that the company, which has not yet been formally named and has no public product, has raised nearly $30 million from investors including 5Y Capital and IDG Capital at a post-money valuation of $200 million. The team has about 10 people. Tian said the lab is building a visual representation system that does not rely on human language. Instead of describing the physical world through natural language, the system uses roughly 200,000 machine-oriented symbols to encode visual information from images and video. He said the model will be trained on about 100 million hours of video data and that a full foundation model is expected in 2027. The report also revisits Tian’s 2024 ByteDance controversy, in which he was accused of interfering with another intern’s model training program and was later dismissed. According to court documents reviewed by Bloomberg, ByteDance sought RMB 8 million in damages, while the court ultimately ordered Tian to pay RMB 500,000 for breach of contract and return the salary he received from the company.

Keyu Tian, the researcher who became widely known in China after a dispute during his internship at ByteDance, has resurfaced as an AI founder.

Bloomberg reported on Oct. 7 that Tian has started an AI lab focused on world models. The company has not been formally named, has no public product, and has a team of about 10 people. Even so, it has raised nearly $30 million from investors including 5Y Capital and IDG Capital at a post-money valuation of $200 million.

A visual system built without human language

Tian said he wants to build an AI visual representation system that does not depend on human language, with the model learning about the real world from roughly 100 million hours of video. That effort places his lab in the same broad field as world-model research pursued by figures including Fei-Fei Li and Yann LeCun.

His argument is that human language, as used by large language models today, is not the most efficient way to describe the physical world. To fully describe a video clip in text, one would need to capture not only people and objects, but also position, motion paths, lighting changes, and physical interactions. In his view, natural language cannot preserve all of that information well enough.

So instead of teaching AI to describe video in human language, Tian said his team is creating a symbol system designed specifically for machine processing of visual data.

According to Tian, the team has already built a dictionary-like visual representation system containing about 200,000 symbols that humans cannot directly read. He compared them to the tokens used in large language models, except these symbols are meant for visual information in images and video rather than text. The goal is to represent, compress, and predict video content in a form better suited to computation.

Tian said the team first needs to give AI its own language, then train the model on about 100 million hours of video.

The work extends his earlier research in visual autoregressive modeling. The broader aim is to train visual models with scalable representation and prediction mechanisms, in a way that resembles how large language models are trained.

He also said the method can already reduce the cost of generating one second of video by at least an order of magnitude, or to one-tenth of the original cost or lower. Bloomberg noted, and the source article repeated, that this claimed efficiency gain still needs to be validated through public technical results.

The team plans to introduce the model’s capabilities through in-person events and public demonstrations first. A full foundation model is expected in 2027. Tian said consumer-facing products are not the immediate priority, and neither is the business model. For now, the focus is on the underlying technology.

The ByteDance dispute and the court ruling

In 2024, while interning at ByteDance, Tian was accused of modifying or interfering with model training programs run by other researchers and was later fired. The incident spread quickly through China’s AI community, and online rumors claimed he had disrupted training for a large foundation model, affected more than 8,000 GPUs, and caused losses worth tens of millions of dollars.

ByteDance later said publicly that some of those online claims were seriously exaggerated or misleading.

In his interview with Bloomberg, Tian acknowledged that he had shut down another intern’s program. He said he believed that person was using more than their share of computing resources and that he wanted to free up compute for other research work.

He described the episode as using an improper method to deal with an improper problem, and said that if he had the chance to do it again, he should have handled it more wisely.

ByteDance later sought RMB 8 million in damages from Tian. According to court documents reviewed by Bloomberg, the court ultimately ordered him to pay RMB 500,000 for breach of contract and to return the salary he had received from ByteDance.

NeurIPS recognition after leaving ByteDance

Despite the controversy, Tian has continued to produce research recognized by the international academic community. In 2024, as first author, he published "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction" and won the NeurIPS 2024 Best Paper Award.

The paper introduced a visual autoregressive modeling method called VAR. Instead of predicting image tokens one by one in the conventional way, the method generates image content progressively from low resolution to high resolution across different scales.

This "Next-Scale Prediction" approach brings visual generation closer to the autoregressive framework used in large language models, while also improving efficiency and scalability in image generation.

Tian’s current startup effort is an attempt to extend that line of work from still images to video and physical-world modeling. His stated end goal is not simply another AI video generation tool, but a foundation model that can understand space, object motion, and physical rules for use in robotics, autonomous driving, and interactive virtual environments.

World models are drawing broader industry attention

World models are becoming a major area of focus outside large language models. Unlike language models, which mainly handle text, code, and reasoning over knowledge, world models aim to give AI an internal representation of the external environment so it can understand how objects move, how they affect one another, and what different actions may lead to.

One of the better-known teams in this area is World Labs, founded by Fei-Fei Li and focused on spatial intelligence. The ABMedia report said AMD announced on Sept. 28 that it would acquire World Labs in an all-stock deal worth about $8.2 billion, with the goal of strengthening its position in robotics, physical AI, and next-generation computing infrastructure through the company’s technology and research talent. The deal remains subject to regulatory review and is expected to close by the end of 2026.

Elsewhere, former Meta chief AI scientist Yann LeCun is also pushing world-model research, while Google DeepMind’s Genie series is exploring real-time, interactive virtual environments.

For Tian’s team, which has only about 10 people and no public product yet, the nearly $30 million in early funding reflects investor interest in both the research track record and the technical direction. Whether its 200,000-symbol visual system and large-scale video training can become a foundation model that genuinely understands the physical world and supports real tasks will depend on future public demonstrations and technical testing.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.