Stanford team introduces OpenWAM to isolate what really improves world-action models

Stanford team introduces OpenWAM to isolate what really improves world-action models

N
News Editor
2026-10-08 23:25:29
A Stanford University team including Fei-Fei Li, Jiajun Wu and Ehsan Adeli has released OpenWAM, an open framework designed to break down the sources of performance gains in world-action models, or WAMs. The paper arrives as interest in the field accelerates, days after Black Forest Labs said its open-source 7B FLUX 3 Action model outperformed the previous top open model, Cosmos 3 Nano, on Nvidia’s RoboLab-120 benchmark with less than half the parameters. OpenWAM starts from Alibaba’s Wan2.2-5B video model, converts it into a causal backbone, and continues pretraining on about 3.34 million robot and human interaction videos totaling roughly 14,600 hours. The framework then combines a 5B video expert and a 2B action expert in a Mixture-of-Transformers architecture and compares four interaction programs under the same backbone, tokenization and training objective. The paper also separates inverse and forward dynamics models into reusable local components and trains them with counterfactual data. In experiments, the causal robot-pretrained backbone lifted VTA success on LIBERO-Long from 68.4% to 97.8%. In transfer tests, a local IDM trained on demonstrations plus counterfactual data reached 84.0% average success, versus 47.0% for a full-context IDM and 21.5% for a local IDM trained only on demonstrations.

Black Forest Labs on Sept. 22 released FLUX 3 Action, an open-source 7B world-action model. In its official blog, the company said the model surpassed Cosmos 3 Nano, previously the strongest open model, on Nvidia’s RoboLab-120 benchmark while using less than half as many parameters. The move drew attention because Black Forest Labs is best known for the FLUX image model, yet it is now pushing into robot control with a system built on video-model priors.

Stanford team introduces OpenWAM to isolate what really improves world-action models 2

This week, a Stanford University team led by Fei-Fei Li, Jiajun Wu and Ehsan Adeli published a paper that tries to answer a more basic question: what exactly is driving the gains in world-action models, or WAMs? Their answer is OpenWAM, an open framework built to separate variables that are usually changed all at once.

Why the field needs a common test bed

Since Nvidia explicitly used the term world-action model in its DreamZero paper in February, related papers have appeared quickly. The problem is that most systems change the video backbone, the training data, the interaction design between video and action, and the inference pipeline at the same time. Scores can be compared. Causes usually cannot.

Traditional vision-language-action models act more like direct policy systems: they take visual input and instructions, then output actions. WAMs add an internal rollout step. They predict what the next few seconds of the scene may look like and what actions would produce that future. Because video naturally captures falling, collision and pushing, large-scale video pretraining can inject a form of physical prior.

That creates a comparison problem. A model can imagine first and act later, act first and predict consequences later, generate both jointly, or keep the two streams separate. If each paper also uses a different backbone and a different dataset, it becomes hard to tell which design choice matters.

Stanford team introduces OpenWAM to isolate what really improves world-action models 3

The Stanford paper points to a second issue. Robot datasets are dominated by successful demonstrations. They show the right way to complete a task, but rarely show what happens if the arm reaches 2 centimeters too far or the gripper closes at the wrong moment. A model may memorize successful trajectories without learning the local relationship between action and physical outcome.

OpenWAM addresses the first issue with a shared backbone and switchable interaction programs. It addresses the second with large-scale counterfactual data.

Paper title: OpenWAM: An Open Framework for Composable World-Action Models

Paper: https://arxiv.org/abs/2610.07922

Project page: https://openwam.stanford.edu

Stanford team introduces OpenWAM to isolate what really improves world-action models 4

Code: https://github.com/OpenWAM/OpenWAM

A causal video backbone built from Wan2.2-5B

OpenWAM starts from Alibaba’s open-source video model Wan2.2-5B. The team continued pretraining it on about 3.34 million robot and human interaction videos, with a nominal duration of roughly 14,600 hours. This stage used no action labels and ran for 14 days on 32 B200 GPUs.

The key modification is causality. Each video block can only attend to the current and previous content, not future frames. The paper contrasts this with the original setup: instead of reading the whole script in advance, the model has to proceed frame by frame in temporal order, which better matches how a robot acts in the world.

The team then used a Mixture-of-Transformers, or MoT, architecture to combine a 5B video expert with a 2B action expert. Each branch keeps its own normalization, projection and feed-forward layers, and the two exchange information only through one joint attention layer.

On top of that shared skeleton, the paper defines four interaction programs:

Stanford team introduces OpenWAM to isolate what really improves world-action models 5

  • VTA, which imagines future video first and then generates actions;
  • ATV, which reverses the order;
  • Joint, which generates video and action together with mutual reference;
  • Decoupled, where the future portions of the two streams cannot see each other.

All four use the same backbone, tokenization and flow-matching objective. The only changes are generation order and cross-modal attention. That gives the comparison a controlled setting that is often missing in this line of work.

Training a reusable action translator

One of the paper’s more distinctive ideas is to separate the inverse dynamics model, or IDM, and the forward dynamics model, or FDM, into independently trainable components.

The IDM takes the current observation, the robot state and an imagined future video segment, then outputs the action sequence needed to realize that future. The FDM does the reverse and predicts how the scene will change given an action.

Both are designed to be local. They do not receive task instructions or long-range history. They only model the relationship between a short visual change and the action tied to it. That makes them less task-specific and easier to combine with any compatible video predictor: the video model decides what should happen next, and the IDM decides how to make it happen.

Stanford team introduces OpenWAM to isolate what really improves world-action models 6

To widen the coverage of these local dynamics models, the team built the LIBERO-Long-CF dataset. They restored simulator states from demonstrations and then executed modified actions, including stopping, reversing, adding noise, and changing gripper timing.

The result was 32,000 segments and about 4.1 million control records, or 29.7 times the size of the original demonstrations. In 72.3% of the segments, object placement changed. Many branches did not complete the task, but that was the point: the model gets to observe consequences of off-trajectory behavior.

Results on LIBERO and real robots

Across the four LIBERO subsets, all four interaction programs posted high success rates. VTA reached an average of 98.6%, slightly above the reported numbers cited in the paper for LingBot-VA at 98.5% and π0.5 at 96.9%. The paper also notes that LIBERO is close to saturation, so the more revealing comparisons come from harder settings.

On LIBERO-Long, Decoupled reached 97.0%, nearly matching Joint at 96.6%. Under a shared backbone, the gap between interaction programs appears smaller than headline benchmark tables might suggest.

For real-world evaluation, the team used two Franka FR3 robot arms on three tasks: making toast, solving the last layer of a 2×2 Rubik’s Cube, and sorting cups by color. VTA and Joint achieved average success rates of 92.1% and 91.9%, respectively.

Stanford team introduces OpenWAM to isolate what really improves world-action models 7

An ablation study showed how much the backbone matters. Holding other conditions fixed, VTA scored 68.4% on LIBERO-Long when initialized from the original Wan2.2. Replacing it with the causal backbone pretrained on robot video lifted the score to 97.8%, a gain of 29.4 percentage points. The authors say that jump reflects the combined effect of the data and the causal redesign.

Transfer tests point to the value of local dynamics plus counterfactuals

The paper’s most important result may be the component transfer experiment. The team froze an IDM trained on LIBERO-Long and paired it with video predictors adapted to four new LIBERO-90 tasks.

A local IDM trained on demonstrations plus counterfactual data reached an average success rate of 84.0%. A full-context IDM reached 47.0%. A local IDM trained only on demonstrations reached 21.5%.

That result suggests two conditions matter at the same time: the model has to focus on local action-outcome mappings, and it has to see enough diverse consequences during training. The paper also makes clear that this is not zero-shot transfer, because the target-task video predictor was fine-tuned on demonstrations with action labels.

The FDM results follow the same pattern. Starting from the same state, the team executed 16 different actions and asked the model to identify which real outcome matched its predicted video. Random guessing would score about 6%. An FDM trained only on demonstrations reached 21.1%. Adding counterfactual data pushed that to 71.3%.

Stanford team introduces OpenWAM to isolate what really improves world-action models 8

What OpenWAM changes

The authors also state the limits of the current work. Component reuse has been validated mainly in simulation, where generating deliberate mistakes is practical. Doing the same on real robots is much more expensive. The FDM also covers only short-horizon prediction.

Even so, the value of OpenWAM is not limited to a few extra benchmark points. In a field where new WAM papers are appearing rapidly, a shared backbone and one-variable-at-a-time test bed can shift attention from who scored higher to why the score improved. The reusable action translator points to another possibility as well: future robot systems may not need to be trained as one monolithic black box. The module that imagines the future and the module that turns that future into actions could be improved separately and recombined as needed.

The team has open-sourced the training, evaluation and composition code, along with pretrained video model weights.

The article was originally published by the WeChat account Machine Heart (ID: almosthuman2014), written by Machine Heart, which focuses on AI.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.