Microsoft Research Asia open-sources Agent Lightning v1.0 for RL training in live deployment environments

Microsoft Research Asia open-sources Agent Lightning v1.0 for RL training in live deployment environments

N
News Editor
2026-10-07 16:09:16
Microsoft Research Asia has released Agent Lightning v1.0 as an open-source, lightweight reinforcement learning framework for AI agents. The project uses a training paradigm called "Harnessed Agentic RL," which lets developers train agents directly in real deployment environments instead of rebuilding agent logic inside a separate training stack. According to the release, the framework is compact at roughly 3,500 lines of code and comes with native Kubernetes support, allowing agents to run as standard Kubernetes jobs without relying on paid commercial sandbox services. The research team also presented an end-to-end coding agent training example built on the Qwen3.5-9B model. Using about 6,000 training samples, the setup improved Pass@1 on the SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute gain of 14.6 percentage points. The source cited for the announcement was Microsoft Research.

Microsoft Research Asia has released Agent Lightning v1.0 as an open-source framework for reinforcement learning with AI agents. The project is positioned as a lightweight system that allows developers to train agents directly in real deployment environments, without reimplementing agent logic inside the training framework.

Training approach and framework size

According to Microsoft Research Asia, Agent Lightning v1.0 uses a training paradigm called "Harnessed Agentic RL." The framework is designed to work with agents already running in production-style environments. The full project contains about 3,500 lines of code.

Native Kubernetes support

The framework includes native support for Kubernetes. Agents can run as standard Kubernetes jobs, removing the need for paid commercial sandbox services.

Coding agent example

The research team also showed an end-to-end coding agent training example based on the Qwen3.5-9B model. With about 6,000 training samples, the setup raised Pass@1 on the SWE-bench Verified benchmark from 41.8% to 56.4%, an absolute increase of 14.6 percentage points.

The announcement cited Microsoft Research as the source.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.