10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention

N
News Editor
2026-08-26 04:11:13
A MarsBit article, republished from the WeChat account QbitAI and written by Noah, has spotlighted a 10-minute embodied AI demo shot in a single take with no visible edits. The video, as described in the report, shows robots from Unitree and Zhiyuan — two different hardware platforms with different architectures, degrees of freedom and sensor systems — operating under what the author calls a shared “brain.” The report walks through a series of scenes in detail: a robot cleaning a window and extending its arm outside after inferring the dirt was on the exterior side; household sorting tasks completed in a cramped roughly 15-square-meter apartment; interruption by an alarm followed by a return to the prior task; and a cross-platform cooperation sequence in which one robot places a scarf around the other so both can carry more items. Another sequence shows a Unitree G1 trying to reach a higher cabinet shelf, then moving a box and attempting to use it as a step. The article also cites unnamed “informed sources” saying the model breaks from prevailing VLA, WAM and traditional world-model approaches. It claims the system was trained on only dozens of hours of video data, though the model’s developer has not been identified and the technical claims remain those of the original author and cited source.

A 10-minute embodied AI demo shot in a single take is drawing heavy interest after MarsBit published a repackaged article from the WeChat account QbitAI, written by Noah. The report says the video contains no cuts, no off-camera teleoperation and no manual instructions, while showing robots from Unitree and Zhiyuan working in the same scene under what the author describes as a shared “brain.”

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 2

The article frames the demo against a broader pattern in embodied AI, arguing that many current showcases are polished to the point of resembling advertisements. By contrast, this video is presented as rougher but continuous, with the author saying the robots can extend an arm outside a window to clean glass, move a box when they cannot reach a high place, and resume a long task after being interrupted.

The report puts special emphasis on the appearance of both Unitree and Zhiyuan robots in the same sequence. According to the article, the two machines differ in hardware architecture, degrees of freedom and sensing systems, yet appear to share one underlying model and cooperate closely. The author calls this a rare example of a general-purpose cross-embodiment “brain.”

A cramped 15-square-meter rental apartment as the test scene

According to the article, the demo was filmed inside a rental apartment of roughly 15 square meters. The room is described as densely furnished with narrow walkways and limited turning space. The author argues that such a setup leaves little room for hidden teleoperators and makes pre-scripted execution difficult.

In the video as described, the two robots perform household chores in parallel within the same space, sensing boundaries in real time and planning their paths. The report says the full sequence shows no collisions, no getting lost and no visible stalls. One damaged robot is even tethered by a lead, yet still appears highly flexible in movement. The article repeatedly stresses that there are no camera cuts, no reshoots and no human instruction during the run.

Window cleaning as a test of force control, perception and reasoning

The first major task in the article’s breakdown is window cleaning. One Unitree robot uses a squeegee on the glass. The author notes that using a squeegee can be harder than using a cloth because too little force leaves the surface dirty while too much force can produce noise, making the task sensitive to both pressure control and spatial awareness.

At one point, after several passes, the robot appears to infer that a remaining dirty patch is on the outer side of the window. It then turns, leans back, extends its head and stretches its arm outward through the window opening to continue cleaning outside. The article argues that this move combines torque control, environmental perception and dynamics calculation, noting that the robot’s arm does not strike the window frame.

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 3

The author presents the sequence as a unified display of vision, touch and dynamics rather than a single isolated capability. The report goes further, describing it as the first time the writer had seen a model make a human-like self-directed inference in a live task by changing strategy when the initial approach failed.

Closed-loop execution across longer task chains

After the window is cleaned, the Unitree robot places the squeegee back in its original location. The report highlights this not as a minor ending motion, but as proof of a complete loop: taking a tool, doing the work and putting the tool away.

Beside it, a Zhiyuan robot approaches a washing machine, removes a cleaned seat cushion and places it behind a sofa, then returns to take out a cleaned doll and places it on the second shelf of a wardrobe. The article says each item is returned to where it belongs.

The report also mentions that the Zhiyuan robot accidentally bumps into glass. The author reads this as evidence of autonomous decision-making rather than remote human control, arguing that camera perception can struggle with reflections on glass while a human teleoperator might have avoided the issue more easily.

For the author, these sequences suggest that the model keeps actions smooth and accurate over long task durations instead of accumulating more errors as the number of steps increases. That point is directly contrasted with what the article describes as a common weakness of conventional VLA models in long-horizon tasks.

Interruptions, then a return to the original task

The next detail in the article is an alarm sound that appears in the video. The author says its purpose was not obvious at first, but after repeated viewing, concluded that the alarm was likely a preset reminder designed to interrupt both robots and redirect them to a special tabletop-and-refrigerator tidying task.

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 4

What stands out in the report is what happens next: once the tabletop and refrigerator work is done, the robots resume the tasks they had been performing before the interruption. The article takes this as evidence that the model can pause work, preserve context and continue from the break point.

The video then moves into a longer home-organization sequence. As described in the report, the Unitree robot picks up a storage bag and places it on the table. The Zhiyuan robot sees food left on top of the refrigerator, judges that it should be refrigerated, and acts on that judgment. It also takes out food that should be thawed. The author adds a personal aside that the mysterious team may be hinting at future household cooking use cases.

The report says there are no step-by-step human prompts visible through this sequence. Instead, the robots appear to decompose the task on their own, plan actions and complete the process from fetching items to putting them away. The author characterizes this as human-level task prioritization and dynamic scheduling.

A cross-embodiment cooperation scene between Unitree and Zhiyuan

The article treats the tabletop cleanup stage as the emotional peak of the video. In the author’s account, the Unitree robot keeps loading objects onto its own body in order to carry more in one trip, until both hands are occupied and it can no longer manage additional items.

At that point, the Zhiyuan robot approaches, picks up a scarf from the table and gently hangs it around the Unitree robot’s neck. Because the scarf is long, Zhiyuan then uses both hands to hold it up, while Unitree lowers its body as if understanding the intention. Zhiyuan passes the scarf over Unitree’s head and leaves it hanging around the neck so more objects can be carried.

The author argues that this was not a fixed division of labor scripted in advance, but rather a cooperation pattern that the model discovered autonomously across two different embodiments. In the article’s framing, the two robots identify their own capability boundaries and complement each other in a way that approaches human teamwork.

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 5

The report even suggests this may be the world’s first cross-embodiment interaction and perhaps the first appearance of a model that can generalize across embodiments. Those are the author’s conclusions from the video. The identity of the model’s developer is still undisclosed in the source text.

Towel-over-shoulder and slipper-by-strap details

After the Unitree robot walks away carrying multiple items, the Zhiyuan robot continues sorting what remains. The article says it picks up a towel and tries to hang it on itself. It fails once, twice and three times before finally resting the towel on its shoulder.

The author interprets the scene as a sign that the robot has found a lower-effort strategy: carrying the towel on a shoulder is easier than holding it in a hand. The repeated attempts before success are presented as evidence that the model can adjust its policy after failure and keep optimizing until it works.

A similar interpretation is applied to the way the robot handles slippers. The report says it lifts a new slipper by its hanging strap rather than by grabbing the slipper body. For the writer, this is another example of an easier and more efficient manipulation choice. The article then links that behavior to self-understanding of the robot’s own embodiment, which it suggests may be central to cross-platform transfer.

Trying to use a box to reach a higher shelf

The article says the most startling sequence comes when the Unitree G1 robot tries to put the scarf onto the third shelf of a cabinet. Because the robot is only 1.3 meters tall, it cannot reach the target height directly.

Instead of freezing or giving up, the robot appears to look for a box nearby and attempt to move it next to the cabinet as a step. It first pushes the box down, then bends in an effort to lift it, only to find that it cannot bend far enough.

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 6

After several unsuccessful attempts, the robot changes tactics and kicks the box toward the cabinet. The article breaks this into three layers of ability: self-directed reasoning and decision-making about how to reach a higher place; exploration of its own physical limits; and spontaneous strategy change, including the use of a foot to manipulate the environment in the absence of explicit teaching.

The author treats autonomous tool use as a hallmark of human behavior and argues that this scene points to a combined understanding of physical rules, self-capability and strategy evolution.

From individual ability to coordinated group behavior

Later in the video, the Unitree robot has kicked the box toward the cabinet but has not aligned it properly. While it is still trying to resolve the problem, Zhiyuan Expedition A3, described in the article as 1.7 meters tall, walks over. The report says it sets down a small cart, switches tasks on its own and chooses to help its counterpart first.

The sequence that follows is described in careful detail. Unitree lowers its head and torso, and Zhiyuan removes the scarf from around its neck and places it on the shelf. The author says Zhiyuan routes the scarf around the back of Unitree’s head before removing it, which the article treats as evidence of fine force control and tight spatial awareness. Zhiyuan then folds the scarf into three parts before placing it into the cabinet, a move the writer reads as an autonomous judgment about better storage.

After helping Unitree, Zhiyuan leaves and places one last piece of clothing into the washing machine. The author guesses that the robot judged the garment draped over the machine to be dirty, though that remains the writer’s inference from the scene.

With that, the 10-minute one-take video ends. The article says it contains a striking concentration of capabilities: cross-embodiment cooperation, autonomous evolution, human-like path selection, strong spatial perception and force control, interruption and task recovery, and the use of tools.

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 7

The article’s reading of the underlying technical route

Beyond the scene-by-scene breakdown, the report cites unnamed “informed sources” saying the most important point is that the model escapes the fixed frameworks of mainstream embodied AI. The article groups current approaches into VLA, WAM and traditional world models, then outlines what it sees as their bottlenecks.

In the piece, VLA is portrayed as a “data intuition” approach that relies on large volumes of vision, language and action data for end-to-end fitting. But the article argues that it remains, at its core, a system that guesses actions from images, lacks physical reasoning, accumulates errors over long task chains, generalizes poorly to unfamiliar scenes and is tightly bound to specific embodiments.

WAM and traditional world models are described as approaches that try to model rules of the world first and plan action later. The article says they suffer from a split between understanding and execution: they may parse scenes and forecast states, but in settings that demand physical dynamics reasoning, they still depend on training data extrapolation and struggle with the variability of real environments.

The author’s broader conclusion is that these mainstream routes share a common limitation: they are forms of data-driven task fitting. In this framing, actions are produced as probability guesses based on large datasets rather than on deep understanding of physical rules, which in turn caps performance at the boundary of the training data.

Claims on data efficiency, scaling and three technical pillars

The article also claims that the model in the video was trained on only dozens of hours of video data. Based on that assertion, the author questions whether scaling-law assumptions should still be taken for granted if such performance can emerge from a relatively small amount of training material.

The report says the model displays at least four capabilities that current VLA and world-model systems cannot match at this stage:

10-Minute One-Take Robot Demo Featuring Unitree and Zhiyuan Draws Industry Attention 8

  • Dynamics modeling: trajectories remain dynamically feasible, with the system able to compute compensation torque and feedforward terms while showing strong extrapolative generalization;
  • Self-cognition: instead of reflex-like reactions associated with VLA or sequential statistics associated with WAM, the system appears to reason about which next action is closest to the goal;
  • Adaptivity: it can fill causal-reasoning gaps, deal with long-tail problems and keep changing behavior while acting;
  • Self-evolution: skills grow and strategies evolve during task completion, as if the system can generate learning goals on its own and reason toward them.

To explain those outcomes, the article proposes three technical pillars. The first is physics-constrained dynamics learning. Unlike VLA and WAM, which the article says are trained in a purely data-driven way, this route allegedly puts physical rules first and uses dynamics prediction as the driver. The writer argues that lifting slippers by the strap, placing a towel over a shoulder and stepping onto a box are not behaviors copied from data, but optimal solutions discovered through interaction with the environment.

The second is unified modeling across embodiments. According to the article, a shared action-representation space and embodiment-adaptation mechanism allow one model to fit hardware platforms from different brands and with different architectures. The report calls this a decoupling of software and hardware in embodied intelligence.

The third is a robust long-horizon closed loop. The article says a global task-scheduling framework allows the model to handle unexpected interference, execution error and task switching on its own, preventing the sort of collapse that conventional systems face when errors accumulate in long workflows.

The team behind the model is still unknown

The article ends by saying it is still unclear which team built the model. The author argues that while many other teams are still showing short, heavily polished, single-task demos, this video presents 10 minutes of unedited, cross-brand, multi-agent, end-to-end task execution with apparent self-evolution in a real setting.

The piece goes on to describe the demo as something that could overturn existing technical routes and move embodied AI closer to its “ChatGPT moment.” Those judgments belong to the original author and the unnamed source cited in the article. Based on the provided text, the developer’s identity, the technical implementation and any independent validation remain undisclosed.

The source article was originally published on the WeChat account QbitAI, authored by Noah.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
20

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.