World Labs unveils Atlas, saying three iPhones are enough to create a Matrix-style bullet-time shot

World Labs unveils Atlas, saying three iPhones are enough to create a Matrix-style bullet-time shot

N
News Editor
2026-09-05 10:53:04
World Labs has introduced Atlas, its new world model built to handle generation, reconstruction and simulation in a single system. The company says the model can use as few as three cameras, or even three iPhones mounted on tripods, to recreate moving camera views around a scene in a way similar to the "bullet time" effect made famous by The Matrix. Unlike standard AI video models that infer a scene mainly from pixels and prompts, Atlas ties each image to a 3D camera pose, letting the model work with explicit spatial context and predict how the world should look from unseen viewpoints. World Labs positions Atlas as a step toward spatial intelligence by combining 3D reconstruction with generative AI. The team says this approach could sharply reduce the amount of visual data needed for scene capture, cutting some workflows from hundreds of photos to only a few images, and in other examples from roughly 2,000 photos to about 30 to 40. Beyond film, games and visual design, the company says Atlas could be useful in architecture, construction, exhibition design, industrial design and robotics, where lower-cost real-to-simulation workflows remain a major challenge.

World Labs has launched Atlas, a new world model that the team says can generate, reconstruct and simulate scenes at the same time by using camera positions, images and 3D spatial information to understand the world.

One of the clearest demos shown by the company centers on the classic "bullet time" effect from The Matrix. In the past, filming a shot like Neo leaning back to dodge bullets while the camera rotates around him would require hundreds of cameras arranged in a green-screen studio, all capturing the scene from different angles at once. World Labs says Atlas can do this with as few as three cameras. It even says three iPhones on tripods can be enough to capture the footage needed to regenerate different views of a camera moving around the scene.

The team gave other examples as well. A person shooting a basketball, or a strawberry falling into milk and sending liquid into the air, can be recorded from a handful of angles. Atlas can then place a virtual camera in positions where no real camera ever stood and regenerate the frozen-in-time effect of a camera traveling through the scene.

Atlas is built around spatial understanding, not only video generation

World Labs says the main difference between Atlas and current mainstream AI video models is that Atlas treats its input images as data with explicit spatial meaning. A typical generative video model can receive an image, infer the scene from its training and shape the result with a text prompt. Atlas works differently. It binds each image to its corresponding 3D camera pose.

That means the model does not only know that an image shows a room. It also knows where in the room the image was captured and from what angle. With multiple images, this forms what World Labs calls spatial context.

Users can then place a virtual camera at any position and time point in the scene and ask Atlas to predict what the world should look like from there. World Labs calls this mechanism New View Prediction, and the team says it could become one of the basic capabilities behind spatial intelligence.

Putting 3D reconstruction and generative AI into one model

Another major change in Atlas is World Labs' attempt to merge two computer vision fields that have long evolved separately: 3D reconstruction and generative AI.

Traditional 3D reconstruction is designed to faithfully recover the real world, so it usually depends on large image sets and triangulation from many angles. Generative AI is strongest at producing things that do not exist in the captured data. World Labs argues that an AI system that truly understands space needs both.

For that reason, Atlas is pretrained to natively support text, images, video, camera poses and 3D information. At this stage, the 3D component is mainly represented through depth maps, giving the model access both to RGB pixel content and to the spatial structure behind a frame.

Fei-Fei Li said this shows that the long-separated tracks of pixel generation and 3D reconstruction in computer vision are starting to converge. For World Labs, the key is to produce pixels that carry spatial context and remain constrained by geometry, which the company describes as an important step toward spatial intelligence.

From hundreds of photos to only a few

World Labs also says Atlas could change the cost structure of 3D content production. In traditional 3D reconstruction workflows such as NeRF and Gaussian Splatting, one of the biggest bottlenecks is dense capture.

Ben Mildenhall, World Labs co-founder and one of the authors behind NeRF, said that fully scanning a room may require 100, 200 or even 300 photos because areas under tables, between chair legs and behind plant leaves all need enough image coverage. For users without experience, scanning multiple rooms can take one or two hours.

Atlas is meant to cut that requirement from hundreds of photos to only a few. Mildenhall said the team is trying to reduce the capture requirement to about three images, which he described as a 50x to 100x drop in data volume.

The reason, according to the team, is that Atlas can first reconstruct the parts that are supported by real images and then use its generative capabilities to fill in areas never captured by a camera.

In a Stanford campus example shown by World Labs, the model used only three to 25 ground-level photos to generate an aerial view of the campus. In some multi-room scenes that previously needed about 2,000 photos for reconstruction, the team said the input can now be reduced to about 30 to 40 images.

This is where the combination of generative AI and 3D reconstruction matters. If an area has never been photographed, a standard reconstruction pipeline can leave a hole. A generative model can infer what is missing from the surrounding spatial context.

Applications from film and games to architecture and industrial design

World Labs says the first and most direct uses for Atlas are still in content creation. Its previous product, Marble, could already turn images, videos or text into Gaussian Splat 3D worlds. But the company found that many creators were following a simpler path: input one image, generate a full 3D scene, capture screenshots from different angles and leave Marble.

That led the team to conclude that many creators do not necessarily need a full 3D asset. Atlas therefore makes New View Prediction a core capability instead of a side feature.

World Labs says this matters in film, advertising, games and visual design. Many AI image and video models can produce new viewpoints, but the positions of furniture, people and objects often drift from one generation to the next. What the company wants instead is a persistent 3D state, where the identity and position of objects remain stable and the creator simply moves the camera rather than re-sampling the whole scene every time.

World Labs believes the same capability could also extend into architecture, construction, exhibition design and industrial design.

A longer-term push into robotics

World Labs says Atlas has a bigger target than Hollywood: robotics. The company previously acquired the robotics startup Cinex, and one of the core technologies involved is real-to-sim and sim-to-real.

For example, a company that wants to train a robotic arm for cable assembly first needs to reconstruct a real factory environment as a simulation environment and then run large amounts of robot training inside it. The problem is that building these simulation environments is expensive.

Fei-Fei Li said the robotics industry's biggest problem right now is data. Real-world robot operation data is hard to collect, and a model cannot be trained on only one standard case. A wire may bend in different directions. Boxes may come in different sizes and colors. Objects may sit in different places in a scene. That means simulation systems also need extensive domain randomization to create many variations.

World Labs says Atlas, with its sparse reconstruction capability, can reduce the cost of turning the real world into a simulation environment.

Over a longer horizon, the company says it wants Atlas to do more than build simulation environments. It wants Atlas itself to become a neural simulator. If the model can understand what the world will look like after a given action is taken, it could simulate the kinds of edge cases a robot might encounter in real environments.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.