‹ BackNewsworld models

world models

Nine AI startups reportedly hit unicorn status within months as investors price founders before products
Bloomberg: ByteDance is building a real-time spatial video model under Zhang Yiming’s oversight
Westlake Univ
2026-09-07 12:19:10

Code World Model splits simulation from rendering as Tim Sweeney joins the discussion

Researchers from Westlake University AGI Lab and Nanyang Technological University have introduced Code World Model, a world-model framework that puts a language model-driven coding agent in charge of maintaining executable world state while a separate video model handles visual generation. The paper argues that current video world models are good at continuing observations, but that is not the same as running a world with goals, rules, memory, and long-range causal consequences. To bridge executable state and image generation, the team uses a proxy representation that carries frame-level spatial constraints and is rendered into proxy video before being fed into the video model together with structured text. The prototype uses MiniMax-H3 as the video backbone with rank-128 LoRA across 50 transformer blocks, totaling about 596 million trainable parameters. Training was run on 8 NVIDIA H800 GPUs for 3 epochs and 3,534 optimization steps. The gameplay dataset includes 157 recordings, about 5.6 hours of source video, and 9,420 training clips sampled every 2 seconds. Each RGB target contains 124 frames at 1344×768 and 24 FPS, with matching proxy sequences at 336×192. During inference, GPT-5.6 Sol acts as the coding agent and GPT Image 2 generates the appearance anchor. In online discussion around the work, a commenter speculated about future Unreal Engine versions, and Epic Games CEO Tim Sweeney replied: 「I don’t know either!」

830
Code World Model splits simulation from rendering as Tim Sweeney joins the discussion
One Year Into AIGC Labeling, Credit Layering Is Emerging as World Models Outpace Old Rules
ACE Robotics chairman says embodied AI could hit a 'ChatGPT moment' by 2027
Inherent unveils Faraday, an AI scientist agent built on a 27B Qwen model
Goldman Sachs says AI is moving into execution, with competition shifting to workflows and world models gaining ground
Unitree Robot
2026-08-20 04:09:00

Unitree founder Wang Xingxing says embodied AI could hit its ChatGPT moment in as little as two to three years

Wang Xingxing, founder of Unitree Robotics, said at the 2026 World Robot Conference that the biggest constraint on embodied intelligence is still weak generalization, and that the industry’s “ChatGPT moment” could arrive in as little as two to three years, or as long as five to ten years. Speaking a day after Unitree’s listing, Wang framed the next major milestone in simple terms: if a robot can be placed in a completely unfamiliar environment and complete roughly 80% of tasks through voice or language instructions alone, embodied AI will have crossed a key threshold. His speech also reviewed Unitree’s 10-year path from quadruped robots to humanoids and outlined a broad product lineup, including the G1 humanoid launched in 2024, the H1 platform, the GD01 mass-produced passenger-carrying transforming mech, the As2-W wheeled-quadruped robot, and the lightweight R1 humanoid. Wang said Unitree has been testing robots in auto factories and in its own facilities, but large-scale rollouts remain limited because robot efficiency and task transfer still lag. He also spent considerable time on data and model training, arguing that humanoid AI needs large volumes of human or internet data for pretraining, combined with real-robot data to align models with the physical world. Wang said the company is also exploring AI-driven self-improving robot development loops, where large models write control code, validate it in simulation, deploy it to physical machines, and refine it through model and human evaluation.

1280
Unitree founder Wang Xingxing says embodied AI could hit its ChatGPT moment in as little as two to three years