Reka, a U.S. AI company founded by former researchers from DeepMind, Google Brain and Meta FAIR, has released Rho-1, a 19-billion-parameter model that can process text, images and video in a single conversation while also producing robot actions.
In the company’s example, Rho-1 can generate an image of a lighthouse, turn that image into motion, change the weather in the resulting video, and then answer follow-up questions about the difference between two video clips. Content produced earlier in the exchange stays in context, so the workflow does not need to be passed to another model.
Video can be edited while generation is still running
Rho-1 also accepts new instructions during video generation. While a clip is being produced, a user can ask to change the weather, the direction of motion or other elements, and later frames will adjust from that point onward.
Reka said the base version runs at about 0.79x real time. Its distilled version reduces the generation process from 99 steps to 8 steps, with the company measuring output at roughly 5.3 seconds of video generated in about 1 second.
Robot control is handled inside the same model
For robotics, Rho-1 predicts what the camera is likely to see next while outputting the robot’s action signals at the same time. That setup removes the need for a separate planning model to decide the robot’s next move. So far, Reka has only shown results on LIBERO simulation tasks and has not released test results from physical robots.
Focus is on continuous interaction
Unified models of this kind have already appeared this year. Nvidia Cosmos 3 can also handle text, images, video, audio and actions in one system. Rho-1, however, is presented with a stronger emphasis on continuous interaction, especially repeated generation, editing and continued reasoning inside the same context.
Still a research preview
Reka said Rho-1 was trained from scratch using 320 H100 GPUs over about three months. The current video resolution is 672×384. The company also noted that long videos can still suffer from structural drift, while video localization and local editing remain unstable.
At this stage, Rho-1 is only available as a research preview. The model weights and API have not been opened.

