Tian Keyu, a former ByteDance intern, has laid out the plan for his new world model startup, saying the company has raised nearly $30 million from investors including 5Y Capital and IDG Capital at a post-money valuation of $200 million. The team currently has about 10 people and has not released a public product.
According to Tian, the company is building a video-focused system around a vocabulary of 200,000 visual symbols. He described those symbols as serving a role similar to tokens in language models, allowing AI systems to process video in a more compute-friendly way. The team plans to train the model on 100 million hours of video and said the approach could reduce generation costs per second of video to one-tenth of current levels or lower.
Tian, the first author of a NeurIPS 2024 best paper, said the startup will first stage a public demo and plans to release a full model in 2027. He also told Bloomberg that he had shut down another intern’s program, saying he did so to free up computing resources for other researchers, while acknowledging that the move was improper. ByteDance fired him in 2024 and later sought 8 million yuan in damages. A court judgment reviewed by Bloomberg ordered Tian to pay 500,000 yuan for breach of contract.
Tian Keyu, a former ByteDance intern, has disclosed the plan behind his world model startup, saying the team has raised nearly $30 million from investors including 5Y Capital and IDG Capital at a post-money valuation of $200 million.
The company has about 10 people at this stage and has not publicly released a product.
A 200,000-symbol system for video
World models are designed to help AI learn how objects move and change in the real world, with potential use cases in video generation and robotics.
Tian was the first author of a NeurIPS 2024 best paper and had previously worked on methods that let AI generate images progressively from coarse to fine detail. His new team is extending that line of work to video by creating a vocabulary of 200,000 visual symbols.
He said those symbols function in a way similar to the tokens used by language models to process text, giving AI a format that is better suited for computation when handling video.
The team plans to train its model on 100 million hours of video. Tian said the approach could cut the generation cost for each second of video to one-tenth of current levels, or even lower.
Product roadmap
According to Tian, the company will first present a public demo, with a full model scheduled for release in 2027.
ByteDance dispute and court ruling
Tian also told Bloomberg that he had shut down another intern’s program. He said he did it to free up computing power for other researchers and acknowledged that the decision was improper.
ByteDance fired him in 2024 and later filed a claim seeking 8 million yuan in damages. Bloomberg, citing a court judgment it reviewed, reported that Tian was ordered to pay 500,000 yuan for breach of contract.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.