Tsinghua team lands in Science Robotics after 22 years of robot soccer work

Tsinghua team lands in Science Robotics after 22 years of robot soccer work

N
News Editor
2026-08-24 07:44:11
A Tsinghua University team led by Professor Zhao Mingguo has published a paper in Science Robotics with ByteDance Seed and China Agricultural University on vision-driven reactive soccer skills for humanoid robots. The study uses an AgiBot humanoid platform for real-robot tests and covers ball finding, chasing, gait adjustment and multi-direction shooting. It also traces a 22-year line from the Tsinghua Vulcan robot soccer team, founded in 2004, to AgiBot’s current hardware platform and the team’s RoboCup work. The paper reports a unified perception-motion reinforcement learning framework, with visual sensing, state estimation, locomotion and physical contact trained together for a continuous soccer task. On the experiment side, the team says the robot can keep reacting when vision is briefly interrupted, and that it achieved stable results in both simulation and hardware tests. The article also notes AgiBot’s Booster T2 platform, unveiled in 2025, and its use in RoboCup and the Robot Competition opening ceremony in 2026.
Tsinghua University’s Professor Zhao Mingguo team has published a paper in Science Robotics with ByteDance Seed and China Agricultural University, presenting a vision-driven reactive soccer framework for humanoid robots. The work uses an AgiBot humanoid platform for real-robot validation and covers ball finding, chasing, gait adjustment, and shots from multiple directions. The paper, titled "Learning Vision-Driven Reactive Soccer Skills for Humanoid Robots," says the robot can complete a full chain of actions using only on-board vision in a dynamic soccer setting. That chain includes spotting the ball, tracking it, adjusting footwork, and shooting. The project sits on a long RoboCup history. In 2004, Zhao led students to found the Tsinghua Vulcan robot soccer team. The following year, the group entered RoboCup with self-built robots. The article says the early machines were about 50 cm tall, moved slowly, and needed human protection against falls. Over the next 20-plus years, the team kept competing, while the robots grew to more than 1 meter tall and learned to stand up on their own, run faster, and shoot harder. The paper also links the academic work to Zhao’s broader research on brain-inspired computing and robot control. His Tianji chip work and an unmanned bicycle project once appeared on the cover of Nature and were named among China’s top 10 scientific advances of 2019. The team’s path also runs into industry. AgiBot founder Cheng Hao once served as the third captain of the Vulcan team. After graduating from Tsinghua, he worked in internet startups and later became vice president of Feishu product at ByteDance. In 2023, he founded AgiBot. Several former Vulcan members joined the startup, and the paper’s first author, Wang Yushi, comes from a newer generation of the same team. Technically, the paper tries to answer a long-running problem: how a humanoid robot can react the moment it sees the ball in a noisy, delayed, and partially blocked real-world environment. The researchers say human players use vision and motion almost at the same time, while robots must do the same with cameras, algorithms, and joint control. Their approach replaces the usual multi-module pipeline with a single perception-motion reinforcement learning framework. To build that system, the team first created a virtual sensing model. Real robots observed the ball at different distances and viewing angles, generating about one hour of data. That data was matched with motion-capture ground truth to measure position noise, detection success rate, update frequency, and latency. In simulation, the ball’s positional noise grows with distance, the camera field of view changes with head pose, and detection can still fail because of motion blur. The vision system updates at about 25 Hz, with average latency of 116 ms. An encoder-decoder structure then helps the policy recover from temporary visual loss. The policy reads the last 50 frames, roughly one second of history, and compresses them into a 64-dimensional hidden state. During training, the decoder tries to reconstruct the ball’s true position and the robot’s dynamics parameters from that hidden state. The actor uses only the observations available on real hardware, while the critic can access full simulation state. Both the decoder and the critic are removed before deployment. The team also added Adversarial Motion Prior training. The dataset includes about 76 seconds of human omni-directional walking and 30 seconds of instep shooting. A discriminator checks how closely the robot’s motion matches the human demonstrations, while reinforcement learning keeps optimizing toward ball pursuit and scoring. Mirror symmetry constraints helped the robot learn left- and right-foot shooting without collapsing into one-sided behavior. The final policy handles ball finding, chasing, foot placement, and shooting in one loop, with the controller outputting joint position commands at 50 Hz. Some behaviors emerged on their own during training. When the ball moved toward the edge of the field of view, the robot would turn its head and body to bring it back into sight. When it could not find the ball near the touchline, it first faced the center of the field and then widened the search. The gait also changed with distance. Far from the ball, the robot used slower steps, with each gait cycle lasting about 0.7 to 0.8 seconds to keep observation stable. As it got closer, the cycle shortened to 0.3 to 0.4 seconds, allowing quicker foot placement. When the goal was behind it, the policy could generate a rotational hook shot: the robot would pivot on its support foot and hook the ball toward the goal with the other leg, avoiding a full re-positioning step. The results were concrete. One second before shooting, raw vision estimates showed a ball-position error of 0.344 meters. After internal state estimation, the error fell to 0.186 meters, a reduction of about 46%. In simulation, if visual detection was interrupted for 0.3 seconds, the robot still kept its ball-contact success rate above 90% and its scoring rate above 50%. Against a stationary ball, the learned policy usually went from start to contact in about 1.5 seconds, compared with 2 to 5 seconds for a rule-based system. On hardware, the robot’s success rate reached 80% to 90% in front-field positions and 60% to 70% in back-field positions, with no falls during testing. Against a rolling ball, the policy still held above 50% success when ball speed stayed below 0.5 meters per second. The same skill module was then folded into the full Tsinghua Vulcan competition stack. In 2025, the team won the RoboCup adult-size humanoid league and the World Humanoid Robot Games, scoring 76 goals and conceding 11 across the two events. In 2026, it defended its title in RoboCup’s Large Size division. The paper’s hardware tests ran on AgiBot’s humanoid platform without special body modifications. Ball data came from the head-mounted camera, and both perception and control ran on the onboard compute unit. The article says that shows mass-produced domestic humanoid robots can now support frontier embodied-AI research, letting researchers focus on algorithms and deploy the trained policy to real hardware. That link from team lineage to hardware platform also shows up in AgiBot’s product line. In July 2025, the company launched its new flagship Booster T2. The Pro version uses an NVIDIA Thor chip and reaches 2,070 TFLOPS of edge compute. Its Booster Studio package combines simulation training, algorithm development, and real-robot deployment in one environment. On August 19, Booster T2 made its debut at the World Robot Conference with a high-explosive motion demo. Three days later, 80 Booster T2 units performed centralized-free autonomous coordination at the opening ceremony of the Robot Competition, spelling out BEIJING, the event logo, and icons for three disciplines on the ground. There was no remote control and no verbal command. The 80 robots’ combined compute reached about 170,000 TFLOPS. From a 50 cm robot that once needed human help to stand, to today’s humanoid platform that can kick on its own and take part in large-scale synchronized performances, the article frames this as a 22-year technical relay between academic work, RoboCup competition, and hardware development.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
10

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.