By: Fangzhou Zhao, Shiyi Cao, Dacheng Li, Charlie Ruan, Sumanth Hegde, and the NovaSky Team

<aside> đź“–

SkyRL-Agent: Deep Research Agent

SkyRL-Agent extends our earlier SkyRL-v0 into an easy-to-use agentic Reinforcement Learning (RL) framework. It is designed to reduce integration overhead by letting users implement tools and verifiers as lightweight single-file modules, where the framework provides built-in ReAct-style agent loops, asynchronous rollout dispatching, and seamless integration with RL training backends.

In this release, we also provide a case study on training a deep research agent using SkyRL-Agent. We walk through how new tools can be added, and share observations from the experiments that surface challenges for complex agentic RL training.

  1. Tools can be resource-consuming, slowing down and destabilizing training. Deep-research tasks often use web-scraping tools with an LLM summarizer. When these tools are under-provisioned, training slows dramatically (e.g., 3.5Ă—) and may hit rollout timeouts. Timeouts produce erroneous rollouts, adding noise to the reward signal, which needs additional gradient updates (~3 steps) to recover from this noise.
  2. Inappropriate tool design can lead to output hacking. With an unrestricted search engine, the model retrieves benchmark answers directly (e.g., from Hugging Face) instead of reasoning on the task. Training with this tool design can falsely encourage the model to search for answers over the internet, which is unavailable during test time.

Model (skyrl-train backend): https://huggingface.co/NovaSky-AI/SkyRL-Agent-WebResearch-8B/

Model (verl backend): [<https://huggingface.co/NovaSky-AI/SkyRL-Agent-WebResearch-8B-verl>](<https://huggingface.co/NovaSky-AI/SkyRL-Agent-WebResearch-8B-verl>)

Code: [<https://github.com/NovaSky-AI/SkyRL/tree/main/skyrl-agent>](<https://github.com/NovaSky-AI/SkyRL/tree/main/skyrl-agent>)

</aside>

Overview

Recent advances in post-training techniques—especially reinforcement learning from verifiable rewards (RLVR)—have sparked a surge of interest in agent training. While there are many reinforcement learning frameworks (e.g., VeRL, OpenRLHF, SLIME, SkyRL-train), each with unique strengths, we are still lacking an agent training framework that can easily leverage any of these frameworks. We argue that such a framework should meet several goals:

  1. Minimal development overhead for users when integrating agents, tools, and environments.
  2. Built-in dispatching of asynchronous calls during rollouts, so users can leverage optimized scheduling strategies (e.g., fixed pool, async batch, or producer–consumer pipelines) without implementing them manually, while still having the flexibility to plug in new dispatching methods if needed.
  3. Flexibility to switch between different training backends to take advantage of their respective benefits.

With these goals in mind, we introduce SkyRL-Agent, an agent training stack built on top of our earlier SkyRL-v0. Users can implement tools and verifiers as lightweight single-file modules, and SkyRL-Agent handles the rest: wiring them into a ReAct-style agent loop and connecting seamlessly to RL training backends such as SkyRL-train or VeRL.

To showcase the framework, we present a case study on training a deep research agent. We walk through how new tools can be added, highlight the training setup, and share our key observations from the experiments. This case study also surfaces several open challenges—such as scaling to more complex tasks and balancing efficiency with reliability—which point toward exciting directions for future work.

Introduction to SkyRL-Agent (Experimental)

SkyRL-Agent is designed primarily for researchers to have a unified interface for implementing agentic tasks. A modular design allows researchers to

  1. Bring in their own tasks with minimum implementation overhead
  2. Use any training backend or simply run the evaluation with the inference-only mode
  3. Modify runtime implementation for a given task (Docker, etc)
  4. Improve the dispatching logic for a batch of trajectories easily
  5. And more …

Abstractions

SkyRL-Agent consists of the following components:

  1. AgentRunner: The main entry point for SkyRL-Agent is the AgentRunner class - it’s responsible for generating trajectories for the given batch of prompts

  2. Trajectory: The trajectory class handles generating a single trajectory for the given instance from the batch. It has three methods:

    1. initialize_trajectory: Set up any runtime environment (e.g., Docker containers and virtual machines) needed for the agent to run.
    2. generate_trajectory: Run the agent loop and get the final conversation and task results.
    3. evaluate_trajectory: Parse the final result and evaluate it with the given task’s verifier.

    The results of both generate_trajectory and evaluate_trajectory are stored in a .result attribute of the trajectory. Each trajectory instance will initialize an Agent instance to generate responses. These three standardized APIs enable any customized agentic task to benefit from any dispatcher strategy, depending on the workload’s characteristics and needs (Figure 1).