Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests

N
News Editor
2026-10-09 03:54:38
Xingdong Jiyuan’s Video Prediction Policy 2, or VPP2, has taken the top spot on the RoboDojo simulation benchmark, posting a 32.26% average success rate and a 39.26 average score. The article says those results put it ahead of GPT-6-Astra, Physical Intelligence’s π0.5 and Nvidia’s GR00T-N1.7. VPP2 also ranked first in generalization, precise manipulation and memory, three capabilities the report frames as critical for robots operating outside tightly controlled settings. The piece goes beyond leaderboard numbers. It says VPP2 was deployed on a real ALOHA bimanual robot and tested on 10 zero-shot manipulation tasks, including grasping, placing, stacking, folding and pouring. There, the model reached a 58.5% average success rate, compared with 40% for π0.5, and recorded the best result in nine of the 10 task categories. According to the report, VPP2 did not rely on extra data or enhancement methods such as Agent RSI. Its core approach was to separate video prediction from action learning and train them in stages. The article also cites additional results on LIBERO-Pro, LIBERO-OOD and long-horizon RoboDojo tasks with a VLM planner, while noting that broader, long-term deployment in real-world settings still needs to be validated.

Xingdong Jiyuan’s world action model VPP2 has climbed to the top of the RoboDojo simulation leaderboard, ranking first on both of the benchmark’s headline metrics. The article reports an average success rate of 32.26% and an average score of 39.26, placing the model ahead of GPT-6-Astra, Physical Intelligence’s π0.5 and Nvidia’s GR00T-N1.7.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 2

The report describes Xingdong Jiyuan as the only embodied intelligence company held by Tsinghua. It says RoboDojo, the benchmark where VPP2 posted those results, is led by the University of Hong Kong’s MMLab and jointly built by nearly 20 top academic institutions worldwide.

RoboDojo is designed to provide a unified and reproducible evaluation standard for embodied intelligence. Its simulation suite includes 42 bimanual manipulation tasks across five dimensions: generalization, precise manipulation, long-horizon tasks, memory and open-vocabulary instruction understanding. The benchmark is meant to test how robots respond when environments, object placements and task combinations change, rather than how well they repeat familiar routines.

Top ranking on RoboDojo’s two core metrics

On the aggregate leaderboard, VPP2 posted an average success rate of 32.26% and an average score of 39.26, both good enough for first place.

The article compares those numbers with GPT-6-Astra’s post-training evaluation on the RoboDojo simulation benchmark. GPT-6-Astra recorded a 22.48% average success rate and a 28.97 average score. By that comparison, VPP2 led by 9.78 percentage points on success rate and by 10.29 points on average score.

The two metrics capture different things. Success rate measures whether a robot completes a task end to end. Average score tracks how far the task progresses even if it is not fully completed. Both metrics aggregate results across the benchmark’s five capability dimensions.

Broken down by category, VPP2 also ranked first in generalization, precise manipulation and memory. The article frames those three as central hurdles for robots moving into real-world settings: whether they can still work after the environment changes, whether they can manipulate objects accurately and whether they can remember earlier steps during multi-stage tasks.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 3

The report also says VPP2 reached those results without extra data and without enhancement methods such as Agent RSI. In its telling, the gains came from pretraining and the base model’s own generalization ability, not from external boosting strategies or score-chasing through larger datasets.

Why RoboDojo matters

The article argues that embodied AI has a measurement problem. Robot demos keep getting more polished and benchmark scores keep rising, yet it remains hard to answer a simple question: which robot is actually smarter and more capable in practical work?

It uses a cup-grasping example to make the point. A robot may perform well on the same tabletop setup it saw during training, then struggle once the cup changes or its position shifts. The problem becomes more obvious in tasks that require several steps in sequence.

That is why RoboDojo draws attention in the piece. Instead of rewarding performance on familiar tasks alone, it pushes robots into harder situations and checks how much capability survives when conditions change.

Real-robot zero-shot test on ALOHA

The article does not stop at simulation. It says Xingdong Jiyuan deployed VPP2 on a real ALOHA bimanual robot and tested it on 10 categories of zero-shot manipulation tasks, including grasping, placing, stacking, folding and pouring.

Zero-shot here means the robot did not receive extra fine-tuning for those evaluation tasks. It had to execute them directly. VPP2 posted a 58.5% average success rate, above π0.5’s 40%, and delivered the best result in nine of the 10 task categories.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 4

In the article’s framing, the leaderboard result shows competitiveness under a standard benchmark, while the real-robot zero-shot result tests whether that capability transfers into a physical environment. Taken together, the two sets of results are presented as the main evidence for VPP2’s technical approach.

Aiming at a long-standing WAM problem

VPP2 follows the world action model, or WAM, route. The basic idea is to use video prediction to understand how the physical world is likely to change next, then convert that prediction into robot actions.

The article says this route has faced a persistent problem for years: a model can generate plausible-looking video without producing a robot that actually performs the task correctly. If a task calls for picking up the cup on the left but the model predicts the cup on the right, the video can still look coherent while the robot goes off target. Another issue is that adding action learning directly into a video model can damage the model’s original generalization ability, leaving the robot more dependent on familiar settings.

VPP2’s answer, according to the report, is that the quality of video prediction sets the ceiling for action performance. Xingdong Jiyuan therefore trained video prediction and action learning in stages rather than mixing them from the start.

A three-stage training strategy

The article lays out a three-stage training process:

  • Stage one: event-level video continued pretraining, so the model learns to predict a full manipulation process.
  • Stage two: fixed-duration video post-training and distillation, so prediction can keep pace with real-time robot execution.
  • Stage three: action expert training, which turns video prediction into concrete actions while trying to preserve the model’s existing generalization ability.

Those stages are meant to deliver two outcomes at once: prediction that generalizes and action that generalizes.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 5

Step one: make video prediction generalize

For the video side, the article says VPP2 is built on Alibaba’s open-source Wan2.1-I2V-14B and trained with a mix of robot manipulation data, human activity data and general video data, covering different robot embodiments and operating styles.

The data pipeline is a major part of the story. Rather than feeding raw videos directly into the model, the team first split manipulation processes into semantically complete segments and paired them with more detailed descriptions. In the article’s example, the instruction is not just “put the cup into the box.” It also specifies which arm, which cup, which box and the full process. The reasoning is straightforward: one vague instruction can map to many possible motion trajectories, and a model that does not know the exact object of the action will struggle to predict the future state accurately.

Another key piece is event-level video prediction training. Traditional short-horizon prediction focuses on the next few frames. VPP2 instead learns the change across an entire operation, from the arm moving toward the cup to grasping it and placing it into the box.

In an instruction-following test on robot manipulation videos, the 14B-parameter VPP2 reached a 90% success rate, while the 64B-parameter Cosmos3 model posted 78%, according to the article. The report uses that comparison to argue that parameter count alone does not decide performance in physical operation prediction.

Step two: turn predictive generalization into action generalization

The article says there are two obstacles here. One is speed. The other is preserving the video model’s generalization ability after action learning is added.

If a robot has to spend several seconds generating video before every move, prediction quality will not be enough. Xingdong Jiyuan’s answer was to speed up the prediction model first. The team converted the event-level predictor into a fixed 8-second video segment predictor, then used consistency distillation to compress multi-step computation into single-step generation. The result, the article says, is about 0.12 seconds of compute time to predict the next 8 seconds of visual change.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 6

On the action side, VPP2 introduces a 0.9B-parameter diffusion Transformer called Action DiT, using a MoT architecture to learn how to generate robot actions from predicted future states.

But the company did not jointly train the video prediction model, Video DiT, and the action expert from scratch. During the early stage of action training, the base parameters of the video model were frozen and adapted only through LoRA. The article says this was done to reduce the risk that action training would damage the model’s existing predictive generalization. In practical terms, the model first learns how the physical world changes, then learns how to convert that understanding into action.

The reported latency numbers are 0.12 seconds for video prediction, about 0.1 seconds for the action expert and roughly 0.22 seconds for total action-segment generation delay.

Additional generalization benchmarks

Beyond the ALOHA real-robot test, the article cites two more generalization benchmarks.

LIBERO-Pro measures manipulation ability when object positions and task requirements change. VPP2 reached a 45.0% overall success rate there, while the best of the other baseline models reached 11.0%.

LIBERO-OOD measures compositional generalization by recombining familiar objects, layouts and task goals into new tasks the model has not seen before. VPP2 posted a 63.9% overall success rate.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 7

Put together, those results support the technical path described in the article: use multi-source data, detailed task descriptions and event-level prediction training to improve physical prediction, then use distillation and staged action learning to convert that predictive ability into manipulation while retaining as much generalization as possible.

VLM planner and long-horizon tasks

The article also argues that real tasks are not only about executing motions. Robots need to decide what to do first and what to do next. Xingdong Jiyuan’s proposed division of labor is simple: GPT handles thinking, VPP2 handles doing.

In the experiments described in the paper, the team used a VLM high-level planner for semantic understanding, memory and task decomposition, then passed explicit subtasks to VPP2 for execution.

On selected long-horizon RoboDojo tasks, adding VLM subtask planning lifted the average success rate from 27.6% to 57.6%, according to the article. The point made there is that strong low-level action capability alone is not enough for complex tasks; coordination between high-level planning and low-level execution also matters.

From GPT to WAM

The article presents this as a possible physical AI pattern: general cognition plus general physical execution. In that framing, general models such as GPT handle intent understanding, reasoning and planning, while WAM systems such as VPP2 predict physical change, generate actions and adjust through execution feedback.

It also says Xingdong Jiyuan is building across the stack, covering the “brain,” the robot body and dexterous hands. The article describes that strategy as “deep full-stack,” arguing that the company is trying to control not only training but also deployment and feedback loops.

Xingdong Jiyuan’s VPP2 tops RoboDojo, leads in simulation and real-robot tests 8

Logistics deployments and current limits

On commercialization, the article says Xingdong Jiyuan has worked with China Post and SF Express and is in regular operation across five provinces and cities and more than 10 logistics centers nationwide.

It notes that logistics environments vary in cargo size, placement and workflow. The same operation can require adaptation when moved from one warehouse to another, which raises the bar for generalization. Whether a model can transfer what it has learned to unfamiliar tasks will directly affect the efficiency and cost of cross-scenario deployment.

At the same time, the article is explicit that existing logistics business progress does not mean VPP2 has already achieved large-scale deployment. Whether the model can run stably over long periods in more real-world settings still needs more validation.

Code and project page are public

The article says VPP2 has been open-sourced and provides the following links:

  • Code: https://github.com/roboterax/video-prediction-policy-2
  • Project page: https://robert-gyj.github.io/video-prediction-policy-2

The original piece was published via the WeChat account Quantum Bit and credited to author Tian Yanlin, then republished by MarsBit.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.