Users on X have spotted a beta feature in the web version of Grok called Imagine Agent Mode. Screenshots suggest the interface is shifting away from a standard chat box and toward an infinite canvas, where the agent can build image and video assets step by step inside a single workspace. xAI has not issued an official announcement, so the scope of the test and any launch plan remain unclear.
An infinite canvas replaces the usual chat interface
According to user screenshots, Grok now shows an “Imagine” entry in the sidebar. Opening it leads to a large canvas area, while the right side displays preset workflow templates such as Create Worlds, Short Film, UGC Product Stories, and Brand Identity. Instead of answering with one output at a time, the agent appears to generate a sequence of assets on the canvas based on a single prompt.
Examples shared by users show a wider production flow: generating and editing multiple images at once, turning still images into video clips, stitching those clips together automatically, trimming footage, applying fade transitions, and exporting the finished result. The change is not just about image quality or video quality. It is about orchestration.
Built for multi-step creative tasks
TestingCatalog said the tool can handle more complex instructions, including prompts such as creating a 1-minute short film, a full comic set, or a UGC product asset package. Another user shared a workflow in which the agent first produced 3 sets of product images and 3 sets of model photos, then combined them into a complete social media content package.
That points to a different operating model from standard text-to-image or text-to-video tools. The system is not only generating outputs; it is also deciding sequence, organizing tasks, and combining results into a usable package inside the same interface.
Extends Grok Imagine beyond standalone generation
The report notes that Grok Imagine launched version 1.0 in February and generated more than 1.245 billion videos within 30 days. It runs on the Aurora engine and supports 720p resolution, 10-second video generation, and native sound. The newly seen Agent Mode appears to add workflow coordination on top of those existing generation capabilities.
This is also described as xAI’s first move to extend the agent architecture from Grok Build’s parallel agents into creative generation. For now, what users are seeing looks like a test build. The feature may be limited to certain regions or an early experiment, and it may not launch in its current form.

