World Labs, the AI startup founded by Fei-Fei Li, has released Atlas, a multimodal world model it says is the first to precisely control camera movement while generating images and video and completing 3D reconstruction. The experience is similar to a panoramic drone, but Atlas does not require shooting every angle. Users input one or a few photos and draw a camera path; the model then fills in unseen walls, object backs, and surrounding space, letting a virtual camera fly along the route. It can generate up to one minute of 1440p video. Unlike ordinary video models, which keep generating pixels from text or images and usually handle camera movement through phrases such as "left pan" or "zoom in," Atlas directly reads the camera's position and angle in 3D space and produces depth plus a full 3D scene. The more input photos, the less the model has to imagine on its own. World Labs says this is why it calls Atlas a "world model": it is not just predicting the next frame, but tracking where the camera is, where objects are, and what another viewpoint should reveal. The company is also using Atlas to turn real spaces into simulated environments for robot training. At present, Atlas is open only to a select group of partners.
Atlas: A Few Photos, a Drawn Route, a Full 3D Scene
World Labs, the AI company founded by Fei-Fei Li, has released Atlas. The company calls it the first multimodal world model that can both finely control the camera in generated images and video and complete 3D reconstruction.
The easiest way to picture it is a panoramic drone, but Atlas never needs you to shoot every angle. Give it one or a few still photos and sketch a camera path; the model will fill in the walls not captured, the backs of objects, and the surrounding space, then let a virtual camera fly along the route. It can output up to one minute of 1440p video.
Camera Pose, Not Pixel Guessing
Mainstream video models keep generating pixels from text or images. Camera moves are usually requested in words like 「left pan」 or 「zoom in」. Atlas does it differently: it reads the camera's own position and angle in 3D space and generates depth plus a full 3D scene behind the frames. More input photos means fewer gaps for the model to fill on its own.
This is exactly why World Labs uses the term 「world model」. Atlas is not just predicting what the next pixel should be; it has to understand where the camera sits, where objects sit, and what a changed viewpoint would reveal.
From Real Spaces to Robot Training
World Labs also applies Atlas to turn real spaces into simulated environments for robot training. For now, access is limited to a select group of partners.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.