Reka unveils 19B-parameter Rho-1 for text, image, video and robot actions

Reka unveils 19B-parameter Rho-1 for text, image, video and robot actions

N
News Editor
2026-10-06 12:56:20
U.S. AI startup Reka, founded by former researchers from DeepMind, Google Brain and Meta FAIR, has introduced Rho-1, a 19-billion-parameter model built to handle text, images, video and robot actions in a single conversation. The company said the model can keep previously generated content in the same context, allowing users to create an image, turn it into video, edit details such as weather, and then compare clips without handing the task off to another model. Rho-1 also supports mid-generation instruction updates, meaning users can change elements like weather or motion direction while a video is still being generated and have later frames adjust accordingly. According to Reka, the base version runs at about 0.79x real time, while a distilled version cuts generation from 99 steps to 8 steps and can produce about 5.3 seconds of video in roughly 1 second. For robotics, the model predicts future camera observations while outputting action signals, removing the need for a separate planning model. Reka said current demonstrations are limited to LIBERO simulation tasks, with no real-world robot test results released. The model remains a research preview, and neither weights nor API access are available yet.

Reka, a U.S. AI company founded by former researchers from DeepMind, Google Brain and Meta FAIR, has released Rho-1, a 19-billion-parameter model that can process text, images and video in a single conversation while also producing robot actions.

In the company’s example, Rho-1 can generate an image of a lighthouse, turn that image into motion, change the weather in the resulting video, and then answer follow-up questions about the difference between two video clips. Content produced earlier in the exchange stays in context, so the workflow does not need to be passed to another model.

Video can be edited while generation is still running

Rho-1 also accepts new instructions during video generation. While a clip is being produced, a user can ask to change the weather, the direction of motion or other elements, and later frames will adjust from that point onward.

Reka said the base version runs at about 0.79x real time. Its distilled version reduces the generation process from 99 steps to 8 steps, with the company measuring output at roughly 5.3 seconds of video generated in about 1 second.

Robot control is handled inside the same model

For robotics, Rho-1 predicts what the camera is likely to see next while outputting the robot’s action signals at the same time. That setup removes the need for a separate planning model to decide the robot’s next move. So far, Reka has only shown results on LIBERO simulation tasks and has not released test results from physical robots.

Focus is on continuous interaction

Unified models of this kind have already appeared this year. Nvidia Cosmos 3 can also handle text, images, video, audio and actions in one system. Rho-1, however, is presented with a stronger emphasis on continuous interaction, especially repeated generation, editing and continued reasoning inside the same context.

Still a research preview

Reka said Rho-1 was trained from scratch using 320 H100 GPUs over about three months. The current video resolution is 672×384. The company also noted that long videos can still suffer from structural drift, while video localization and local editing remain unstable.

At this stage, Rho-1 is only available as a research preview. The model weights and API have not been opened.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.