Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference

N
News Editor
2026-09-29 04:37:11
Tsinghua University, Infinigence AI and Shanghai Jiao Tong University have open-sourced APXInf, an on-device inference engine built for embodied AI models running on robot hardware. The project targets a problem that is becoming harder as robot models iterate faster and commercial deployments move closer to real-world delivery: how to run large embodied models efficiently and reliably on devices with tight limits on compute, memory and power. According to the project description, APXInf cut inference latency for the π0.5 model on NVIDIA Thor in FP8 mode from 278 ms to 26 ms without changing the model itself, a roughly 10.7x end-to-end speedup. That pushed control frequency to 38.46 Hz, placing it in the real-time range for robot control. The team also said the engine reached 92.2% success on Thor FP8, 92.8% on Thor BF16 and 92.0% on Orin BF16 in a LIBERO-10 evaluation, compared with 92.4% for the π0.5 reference implementation. Beyond raw speed, APXInf is positioned as reusable infrastructure. The repository turns model adaptation, optimization and validation into a workflow that can be reused over time and called by agents. The project has already added support for π0-fast, GR00T and Qwen Drive, with AMD and Chinese chip backends listed on the roadmap.

The embodied AI field saw a burst of model releases in the second week of September. On Sept. 8, Songyan Dynamics released HERON-World Model. On Sept. 9, AgiBot unveiled AGILE 2.0 and GE-Act 2.0. On Sept. 10, Unitree open-sourced UnifoLM-WLA-1.0, a 6B-parameter model trained on about 2,500 hours of real-robot data and designed to handle 64 tasks with one model. By Sept. 15, Songyan Dynamics had added HERON-CRA. That made five models from three robot makers in eight days.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 2

At the same time, robots were moving faster into real-world deployment. On Sept. 20, Qiyuan Robotics held a product launch and formally started sales of two personal robots, with the Q1 priced from RMB 19,999. Earlier, on Sept. 10, UBTECH said it had secured overseas orders worth more than RMB 50 million, covering products including Walker C1 and YouWorld U1. As delivery becomes more concrete, the yardsticks used by capital markets are also shifting toward repeat purchase rates after adjustments, operating net cash flow, and actual fulfillment costs, meaning the spending required for deployment, tuning and maintenance after delivery.

Those two trends meet at one bottleneck: how to deploy increasingly capable embodied foundation models onto robot hardware that has very limited compute resources, while keeping performance efficient and stable. That is the problem APXInf is built to address.

APXInf targets speed on-device and a repeatable optimization path

On Sept. 15, Tsinghua University, Infinigence AI and Shanghai Jiao Tong University open-sourced APXInf, an on-device inference engine for embodied models.

The project is framed around two questions. First, in edge environments constrained by compute, memory and power, how can embodied models reach usable inference speed on robot hardware? Second, as models keep evolving, how can developers build an optimization capability that keeps adapting instead of falling behind?

For the first question, the project points to a specific benchmark. Without changing the π0.5 model itself, APXInf said full-stack end-to-end optimization reduced inference latency on NVIDIA Thor in FP8 mode from 278 ms to 26 ms. That is about a 10.7x speedup, lifting control frequency to 38.46 Hz and into the real-time range for robot control.

For the second question, the answer sits in how the repository is organized. APXInf turns model adaptation, optimization and validation, work that often depends on a small number of specialists, into a workflow that can be reused continuously and called by agents.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 3

Project repository: https://github.com/RLinf/APXinf-robo

Why embodied AI needs dedicated edge inference optimization

To see why APXInf exists, the article first places the inference engine inside a robot’s control loop. A typical embodied robot consists of a main controller, a compute box and peripherals. In one control cycle, the main controller collects observations, the compute box runs inference and generates an action chunk, and the result is sent to the arm or chassis for execution before the next frame begins. The inference engine sits in the middle of that loop. It determines how many milliseconds each inference takes, what control frequency is possible, and how large, hot and expensive the compute box needs to be.

That position creates a very specific set of requirements: small-batch inference, real-time response, low latency jitter, and stable invocation by the main controller through websocket or ROS. Those are not the strengths of cloud-oriented inference frameworks.

General-purpose inference stacks usually lower models layer by layer through a unified intermediate representation and then generate executable code in the backend, aiming to cover as many models and hardware targets as possible with one compiler stack. Systems such as vLLM and SGLang are built around cloud throughput. They work well in environments with many model types, large batches and broad scheduling room. Embodied edge deployment has none of those conditions.

Three bottlenecks show up in real deployments

The first is edge performance. On-device systems are constrained at the same time by compute, bandwidth, power and thermals, yet each inference still has to complete multi-view perception, model forward passes and action generation while responding to the main controller at a stable rhythm. The article says memory bandwidth on edge modules can be 4x to 8x lower than on discrete GPUs, while power budgets can differ by 5x to 10x. A model that runs smoothly in the cloud can lose a lot of effectiveness when moved to Thor or Orin.

The second is labor and time. Deploying a model to edge hardware is not a copy-and-run task. It is a full systems project that spans low-level architecture adaptation, core operator compilation, quantization, software and hardware tuning, and simulation-based validation. That process often takes weeks and depends on cross-domain specialists who understand both inference systems and operator optimization. Those people are scarce. Small hardware changes can also wipe out earlier work. Once the chip changes, operator choices, memory layouts and pipeline arrangements may all need to be rebuilt.

The third is stability. A demo that runs once is not the same as a system that can run for long periods. Production deployment has to deal with continuous perception and control, limited resources and coordination across multiple modules. The risk is not only slower performance, but also jitter, freezes and loss of state control.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 4

The article summarizes the workload with a multiplicative formula:

Total work to deploy a model onto a robot body ≈ (one integration + one tuning cycle) × number of robot body models × number of chip platforms × number of model iterations.

Each variable on the right side is growing. Robot body variants are increasing, chip platforms are diversifying, and model iteration cycles are getting shorter. The five models released in eight days at the start of the article are presented as a direct example. In that setting, simply hiring more people is not a realistic answer. What is needed is reusable infrastructure that can push single-run inference close to hardware limits while keeping the next model or next chip from starting at zero.

How APXInf pushes inference into the real-time range

APXInf does not treat a unified general-purpose IR as its first priority. Instead, it builds specialized execution paths for model families. Model structure, weight layout, memory space, operator fusion schemes and execution order are expressed directly in code. Only capabilities that prove reusable across multiple models are later abstracted into shared modules.

That design lets optimization go deeper into model structure and hardware characteristics. It can tune operator selection, memory layout and execution flow for a specific target, cutting the overhead that comes from broad compatibility.

At runtime, APXInf keeps only the control features needed for real-time edge inference:

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 5

  • Operator execution uses CUDA Graph for full-graph capture and steady-state replay, while kernel selection is generated and persisted through Autotune.
  • Memory is held by the model layer in a fixed workspace with fixed shapes, pre-allocation and stable addresses, reducing data movement on the hot path.
  • Scheduling is centered on small-batch real-time inference, without introducing throughput-oriented features such as Continuous Batching and Paged Attention.

The runtime avoids complex heuristics. The tradeoff is a path that is predictable, reproducible and auditable.

The same specialization carries into the build process. During compilation, APXInf queries the local GPU’s compute capability and compiles kernels only for that architecture. Source code for CUDA kernels, CUTLASS and FlashAttention is bundled in the repository, so deployment does not require Docker or external framework dependencies, avoiding container images that can run into multiple gigabytes.

The article uses Orin to illustrate the optimization ladder. Starting from a 1,300 ms baseline, the system moves through torch.compile, Pipeline, Graph, Kernel and Pruning, ending at 119 ms, an improvement of more than 10.9x. Thor shows a similar pattern.

Speed is only part of the story

The official evaluation protocol uses all 10 tasks in LIBERO-10, with 50 episodes per task, a fixed seed of 7, a replan step size of 5, and 500 rollouts in total. Reported results were 92.2% success on Thor FP8, 92.8% on Thor BF16 and 92.0% on Orin BF16. The π0.5 reference implementation posted 92.4%.

APXInf also puts weight on long-run stability. The low-level stack focuses on high-performance operators, while the main body of the inference framework is written in Rust. The article says Rust gives the team low-cost, fine-grained control over the system runtime, which helps with scheduling and concurrency management.

It also points to Rust’s enforced RAII behavior and ownership system as a way to narrow the risk surface for memory issues and support safer, more robust lifecycle management of resources. Unsafe code is tightly confined to fixed FFI boundaries. The system-level value, as described in the article, is to eliminate hidden failures such as wild pointers and data races before deployment. For a robot that has to run continuously, the requirement is not just to hit 38 Hz once, but to still be at 38 Hz after eight hours of operation.

Keeping up as models change

Specialized execution paths improve performance, but they also create a maintenance problem. If each model family needs a hand-built path, then every model update or hardware generation shift means rebuilding that path. If the work still depends on a small number of experts, it will not keep pace with embodied model iteration.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 6

APXInf’s answer is to productize the engineering capability itself. The project turns model integration, pre- and post-processing, robot-body adaptation, performance tuning and deployment validation, work that used to live in scattered expert experience, into an engineering workflow that code agents can understand, call and iterate on.

That changes the division of labor:

  • Agents run the process: reading the PyTorch reference, generating the execution ledger, implementing the model’s static path, checking operators one by one, running regressions, and then completing Autotune, documentation and iteration.
  • Humans set the standards: architecture boundaries and module responsibilities, kernel contracts and safety rules, acceptance thresholds for accuracy, performance and tasks, and decisions on when common capabilities should become shared abstractions.

In the article’s framing, agents are no longer limited to simple assistance. They take on the implementation work that consumes the most expert time, while people move toward judgment and decision-making.

That makes validation a central moat. APXInf describes three layers of protection:

  • Layered end-to-end validation, from single operators and layers to the full model, with cross-checking between eager and graph paths.
  • A fail-closed principle, where unsupported parameters or hardware trigger direct errors instead of silently producing wrong outputs.
  • Isolation by model family, with shared capabilities sinking only to the kernel layer and any change requiring strict layered regression.

What this means for companies and for the sector

For embodied AI companies, APXInf is presented as a way to turn inference optimization from a one-off project into a sustainable engineering capability. When an in-house model changes, teams do not have to wait in line for specialists again. When a chip changes, they do not have to discard all prior tuning experience. Team size is no longer a hard cap on integration speed. Expert know-how can become a reusable team asset instead of remaining personal intuition.

For the broader sector, the project suggests a different way to think about infrastructure. In the past, inference frameworks were often judged by how many models they supported, how much hardware they adapted to, and how complete their operator libraries were. Those metrics assume integration cost is fixed, so broader coverage is better. The article argues that once integration cost itself starts to fall, the more important question becomes how long it takes to bring up a new model or a new robot body. That is a measure of how quickly an embodied AI company can turn a new model into actual production capacity.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 7

From Mizar to APXInf

The article says APXInf is not an experimental project built from scratch. It is described as a direct extension of Infinigence AI’s longer-term work in edge inference. Its technical lineage comes from Mizar, an inference acceleration engine aimed at smart terminals.

During the Mizar phase, the company built capabilities around AI PCs, AI boxes and all-in-one devices, including heterogeneous hardware adaptation, local model deployment, inference acceleration, memory optimization and low-power operation. According to the article, that stack supported several-fold increases in local model scale, doubled inference performance, and delivered an 18% improvement in intelligence level under the same compute resources.

The article also says the route has already passed mass-production testing. In February 2025, Infinigence AI reached a deep partnership with Lenovo to embed the Mizar engine into Lenovo’s new generation of AI PCs. In November of the same year, the two sides reached a pre-installation deal covering more than 10 million AI PCs. APXInf is presented as the next evolution built on that base.

AI PCs and embodied edge systems do not carry the same workloads, but they face the same constraints: usable inference speed under limited compute, memory and power. To handle the added demands of embodied scenarios, including more complex multimodal execution chains in VLA models and millisecond-level real-time control, APXInf extends the stack with deeper low-level optimization and an agent-oriented engineering workflow.

Why it launched inside the RLinf open-source ecosystem

Another part of the strategy is engineering continuity. APXInf was first released in the RLinf open-source ecosystem because model training and on-robot inference are two stages of the same engineering chain.

RLinf covers reinforcement learning training, evaluation and real-robot workflows. But once an embodied model is trained, it still needs an inference engine designed for small batches, low latency and limited power before it can run on the robot body. APXInf fills that gap, extending RLinf from training and evaluation into edge inference and deployment so developers can move through training, validation and on-robot execution within one ecosystem.

Roadmap and open collaboration

As of the article’s publication, APXInf had already completed support for π0-fast, GR00T and Qwen Drive, extending model coverage across robot manipulation and mobile-agent scenarios.

Tsinghua, Infinigence AI and Shanghai Jiao Tong University open-source APXInf for embodied AI inference 8

On the hardware side, AMD and Chinese chip backends are already on the roadmap. The stated goal is to support efficient and stable operation for more embodied models across more edge devices.

The article closes by saying that if embodied intelligence is to move fully into the physical world, models, robot bodies and the infrastructure layer between them are all necessary. Faced with the scale of integration and validation work, APXInf chose an open-source path. Developers deploying embodied models such as π0.5 onto robots, those working with RTX 4090, Jetson Thor or Orin hardware, and contributors interested in adding new models, chips, model-integration skills or the agentic deployment workflow itself are invited to try the project, file issues or submit pull requests.

GitHub: https://github.com/RLinf/APXinf-robo

Quick Start: https://github.com/RLinf/APXinf-robo#build-apxinf-robo

Acknowledgments

The article says APXInf was shaped by several open-source projects and thanks their communities for their contributions. The list includes FasterTransformer, Tensor-LLM, llama.cpp, vLLM, sgLang and FlashRT.

  • FasterTransformer: https://github.com/NVIDIA/FasterTransformer
  • Tensor-LLM: https://github.com/NVIDIA/TensorRT-LLM
  • llama.cpp: https://github.com/ggml-org/llama.cpp
  • vLLM: https://github.com/vllm-project/vllm
  • sgLang: https://github.com/sgl-project/sglang
  • FlashRT: https://github.com/flashrt-project/FlashRT

This article was sourced from the WeChat public account Machine Heart (ID: almosthuman2014), written by Machine Heart.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.