Microsoft Agent Lightning v1.0 Trains AI Agents on Real Deployment Harnesses
Microsoft's Agent Lightning v1.0 trains real agent harnesses directly, lifting Qwen3.5-9B 14.6 points with 6,000 samples and native Kubernetes support.
Summary
Microsoft Research Asia open-sourced Agent Lightning v1.0 on October 7, 2026, rebuilding the agent reinforcement learning framework in about 3,500 lines of code. Its Harnessed Agentic RL paradigm trains the same harness used in deployment instead of recreating its interaction loop inside frameworks such as verl, AReaL or slime. An OpenAI-compatible LLM proxy records prompts, responses and log probabilities while leaving harnesses such as mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex unchanged. The control plane combines an API Gateway, Rollout Controller and verl-based Customized Trainer, with a sample adapter converting completed rollouts into training data.
Agents run locally or as standard Kubernetes jobs across self-managed clusters, cloud Kubernetes and local infrastructure, avoiding paid sandboxes such as Modal Sandbox and E2B. Collocated Async RL shares GPUs between rollouts and model updates, pauses new requests during updates and delivered roughly 2x the end-to-end speed of synchronous RL while using fewer GPUs than conventional asynchronous training. A pipeline using SWE-smith, mini-SWE-agent and Qwen3.5-9B, including data cleaning, environment construction and reward-hacking safeguards, raised SWE-bench Verified Pass@1 from 41.8% to 56.4%, a 14.6 percentage point gain, with about 6,000 open-source training samples and no large-scale compute. Real harnesses still create retokenization and sample-merging errors, sample-level advantage bias, distorted loss normalization and unpredictable backend workloads. Experiments found rollout-level advantage calculation and normalization improved validation reward and stabilized policy entropy.
Positives
- About 3,500 lines of code provide a complete agent reinforcement learning control plane designed for inspection, modification and extension.
- Qwen3.5-9B reached 56.4% Pass@1 on SWE-bench Verified, up 14.6 percentage points from 41.8%.
- About 6,000 open-source samples produced the benchmark gain without large-scale compute.
- Collocated Async RL delivered roughly 2x the end-to-end speed of synchronous training while requiring fewer GPUs than conventional asynchronous RL.
- Native Kubernetes jobs support self-managed clusters, cloud Kubernetes and local infrastructure without Modal Sandbox or E2B.
- An OpenAI-compatible proxy connects existing agent harnesses to training without changing their execution code.
Risks & concerns
- Retokenizing harness text can shift token boundaries, preventing adjacent model calls from merging cleanly into one training sample.
- Subagents and context summarization can split rollouts, causing sample-level advantage calculations to count some rollouts repeatedly.
- Sample-count loss averaging can overweight rollouts that produce more samples because of harness behavior rather than task performance.
- Unknown sample counts and lengths must be scheduled onto fixed GPU and parallelism configurations only after each harness finishes.
- Collocated Async RL temporarily stops accepting new model requests while updates complete and in-progress requests drain.