Scalabot unveils HERON-World Model for action-conditioned prediction across multiple agents
Scalabot has introduced HERON-World Model, which the team describes as its first world model designed to predict what happens next after a robot takes an action. Given a current frame and an action input, the model generates a forecast of how the scene may evolve. The company said the system goes beyond single-agent prediction and can also model coordination, interference, and shared physical outcomes when multiple agents act at the same time. According to Scalabot, HERON-World Model was evaluated offline under the open-source WorldArena 1.0 Track 1 benchmark and posted an EWMScore_P of 75.80. In category scores cited in the release, it recorded 95.60 in physical interaction, 99.20 in spatial perspective, 92.61 in motion dynamics, and 56.66 in trajectory accuracy. The demos highlighted three main capabilities: modeling physical laws, maintaining state memory and temporal consistency, and forecasting multi-agent interaction. The team said the training pipeline combines Ego first-person human operation video, simulation data reconstructed from real scenes, and real-robot data. It also uses a MoT+MoE architecture. In the setup described by Scalabot, MoT is used to preserve the Vision-Language Model’s prior world knowledge while training the generative branch, while MoE splits generation into high-noise and low-noise experts to handle coarse scene structure and fine interaction details separately.








