New-ZZZ
RU / EN
AI Agents 2 October 2026

New RL Method: How to Train AI Agents Across Multiple Environments

N
New-ZZZ desk
X @huggingface · 3 hours ago

Researchers from Hugging Face and the @adithya_s_k team presented an advanced guide on multi-level reinforcement learning (RL), which fundamentally changes the approach to training AI agents.

Instead of directly modifying the testing environment, the authors developed an innovative method using a "proxy." This proxy layer intercepts the exact token IDs and logits (logprobs) generated by the model (e.g., vLLM), regardless of which API the code agents use (OpenAI, Anthropic, Gemini).

This allows agents to be trained on data obtained from multiple testing environments without the need to change the base code. The results showed that the LFM2.5-2.6B model improved from 42% to 54%, and OpenCode — from 34% to 58%.

Furthermore, the authors demonstrated that simple fine-tuning on successful runs of a large model (Qwen3.8-27B) is less effective than a structured RL approach. The key takeaway is that a practical, multi-level approach surpasses simple weight copying.

Why it matters

  • —A practical and open framework for multi-level RL is presented, significantly improving the efficiency of training AI agents.
  • —The method uses a 'proxy' to collect data from various APIs and test environments without modifying the agent's base code.
  • —Results demonstrate that structured training (RL) is much more effective than simple fine-tuning on successful cases.

Key facts

  • The LFM2.5-2.6B model improved from 42% to 54% after training on 4 test environments.
  • Implementing a proxy layer allows capturing precise token IDs and logprobs from various APIs (OpenAI, Anthropic, etc.).
  • Training on multiple environments increased agent accuracy and reduced the number of tool calls by 31%.
  • The authors proved that simulating training based on a large model (Qwen3.8-27B) is less effective than a specialized RL approach.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related