SenseTime launched SenseNova U1 Pro at WAIC 2026, presenting it as a delivery system for complex multimodal tasks rather than a model focused only on generating a single image on command. In the company’s description, the system is built to complete a chain of work around a target outcome: understanding, planning, organizing information, generating multimodal content, checking, revising and delivering the final result.
The release was introduced with an ultra-long 8K image themed around “WAIC Ninth Anniversary 2018—2026.” The picture spreads from left to right as a timeline and includes highlights from each edition of the event. The original report said the full-resolution file was 51 MB and that the text remained clear after zooming in.
The report grouped U1 Pro’s main selling points into three areas. First, native 8K output, with the emphasis not just on higher pixel count but on preserving text, composition and detail across a very large canvas. Second, what the article described as an interleaved text-and-image mode of thinking, where the model can move through sketching, refinement, coloring, checking and adjustment around a defined goal. Third, a stronger focus on final deliverables, with SenseTime targeting scenarios such as infographics, urban planning, film storyboards, academic posters and commercial design. The stated aim is to reduce repeated trial-and-error generation and produce results that can actually be used.
The article framed that shift in practical terms: can a model finish more of the job itself, instead of leaving people to clean up the result afterward?
Four tests used to examine the model’s output
To check how U1 Pro performs in practice, the report put it through four tasks covering panoramic composition, complex poster design, stylized visual execution and commercial-style output.

An 8K panoramic image of the 24 solar terms
The first task asked U1 Pro to generate a horizontally extended 8K image based on the 24 solar terms. This was presented as a test of detail control. The names and order of all 24 terms had to remain correct, each one had to match the proper seasonal imagery and color palette, all 24 vertical sections needed to stay distinct while maintaining a unified visual language, and the transition from spring to winter had to look natural.
According to the article, U1 Pro placed all 24 points within one continuous horizontal layout instead of turning them into 24 separate wallpaper-like panels of equal size. It also varied the height of mountains, plant placement and negative space from one section to another, creating a rhythm that the report compared to a mix of folding screens and a long handscroll.
A laboratory-style visual poster
The second test asked for a visual poster titled “How Machines Observe and Understand Humans.” The prompt required a figure standing with their back to the viewer and slightly turned to one side, with detection boxes, coordinate axes, geometric circles, motion paths, gaze tracking, spatial grids and data nodes layered over the head and upper body.
The report said many image-generation models default to a blue palette when they encounter words such as “technology” or “robotics.” U1 Pro did not follow that path in this example. Instead, it used dark gold lines, a black background and a paper-like grain to build a clearer visual hierarchy. The center of the image remained focused, and while the information density was high, the frame did not collapse into disorder.
A glazed-texture traditional architectural landscape
The third task was a landscape scene with traditional Chinese architectural elements rendered with a glazed texture. The prompt required green-blue mountains, blue rivers, waterfalls, pagodas, gate structures, pavilions, covered bridges, palaces, peach blossom groves, pine trees, auspicious clouds and one futuristic building that still used Eastern structural language.

Material requirements were also specific. The mountains were supposed to look like a fusion of glaze, jade and enamel, with flowing highlights, layered surface texture and golden edging. The water needed transparency and reflection. The buildings could not look like cheap 3D assets.
The report’s conclusion was that U1 Pro managed to organize the glazed mountains, water system, traditional buildings and futuristic architecture into a relatively complete visual world, showing that it could execute against a complicated style brief.
A movie poster meant to look commercially ready
The fourth test moved toward a more commercial use case. The prompt called for an original movie poster with a high level of finish, suitable for theatrical promotion, and with a tone positioned between Eastern poetic aesthetics, suspense epic and modern art cinema. It was also expected to carry strong visual impact and a polished look.
The article said the final image again kept its text readable and delivered the kind of result that felt close to direct commercial use.
Looking across all four tasks, the report argued that U1 Pro’s strength is not limited to 8K resolution. It appears more capable of organizing a complete visual objective in one pass, bringing together information, layout, characters, materials and style inside the same job. The article also said the model could handle structural diagrams and commercially usable imagery.

From one-off generation to checking and revision
The article said the common thread in these tests goes beyond image quality. Whether it is the 24-solar-term panorama, the complex posters or the glazed landscape, U1 Pro has to keep multiple requirements in view at once and continue managing information, composition, style and detail inside a single image.
That maps onto a practical weakness in current image-generation tools. Many models can already take repeated natural-language edits, but once a task becomes more complicated, the result can still slip out of control. Fix one local area and other parts may change with it. An image can look polished at first glance while its text and structure fail closer inspection. After several rounds of edits, textures, figures and backgrounds can drift as well.
SenseTime CEO Xu Li summed that issue up on site with a short line: “Being interactive does not mean being deliverable.”
In the article’s telling, U1 Pro tries to turn one image-generation request into a longer creative process. The “thinking” here is not a visible block of textual reasoning shown to the user. It is closer to a continuous creation process in which text and image work through the task together.

In an exchange with QbitAI, SenseTime co-founder and chief scientist Lin Dahua said the company had already observed an early form of continuous creation during the U1 stage. The model could sketch first, then add detail and color, gradually building a complete image. That led the team to see a path for visual models to move closer to the way designers work and to push image generation into real content-design production.
With U1 Pro, the process is described as a closed loop: understand the goal, plan the task, organize information, generate content, check for problems, keep revising and then deliver the result.
The WAIC anniversary handscroll shown at the start was used as an example. To produce that piece, the model had to digest material covering nine years, decide how events should be distributed, work out the transitions from one year to the next, arrange landscapes and urban scenes, decide where text should go and keep the whole work in a consistent Eastern visual style.
How SenseTime says it handled native 8K generation
Native 8K output was one of the clearest labels attached to U1 Pro at launch, but the report also stressed the technical burden that comes with a much higher resolution. As resolution rises, the number of visual tokens rises as well, which increases both Attention computation and memory demand.
To deal with that, one of SenseTime’s key choices was the use of 32×32 large patches. Lin Dahua said many common image-generation models may use 16×16 patches. If one patch corresponds to one token or one group of tokens, then doubling the edge length cuts the total number of visual tokens to one quarter of the original amount.

That is roughly equivalent to dividing an image into much larger cells. Fewer cells mean lower computational pressure. The trade-off is that larger cells are more likely to lose fine detail.
To address that, the team added adaptive noise control and used more targeted training strategies for detail-heavy regions. The report also said patches retain some overlap, while spatial sampling, loss design and model structure were optimized as well. Together, these methods are meant to control context size, prevent an explosive increase in cost at 8K resolution and preserve small text, texture and structure as much as possible.
The article summarized the approach in simple terms: enlarge the patch size to reduce total load first, then recover internal detail through overlap and more refined training.
The product shift SenseTime is trying to describe
From an industry perspective, the article said the more important point about U1 Pro is the product shift it represents. To explain that shift, it compared the trajectory to AI coding.
In the first stage came Copilot, helping professional developers complete code. Then came what the article called Vibe Coding, where users express needs in natural language. After that, coding agents began to decompose tasks, write code, call tools, test and repair, taking on more of a complete engineering workflow. In that process, the value of the model moved away from how many lines of code it could write and toward whether it could get the project done.

The article argued that multimodal content is starting to follow a similar pattern. The first stage is point generation. The second is intent-driven iteration, where users can keep revising. The third points to system-level content delivery. At that stage, a model is expected not only to generate one image, but also to understand the goal, organize information, maintain long-range consistency, check errors and deliver something usable.
The lesson SenseTime says it drew from coding is that once a technology crosses an industrial threshold and starts changing productivity in a real way, it may open a larger commercial market. On that basis, the company believes a text-image interleaved workflow and a unified understanding-generation system could also open a differentiated track in visual design.
The report did not present U1 Pro as a finished answer. It said asynchronous generation means users trade waiting time for higher completion quality. Editing ability, cost and stability still need validation in real projects. It also noted that professional design work does not depend only on one final image; layers, vector elements, version management and team collaboration still matter.
The article ended on a directional point: once image generation moves away from repeated lottery-like prompting and starts taking responsibility for the result, competition in multimodal agents truly begins.

