QwenGyre: An Elastic RL Framework for Training xLong-Horizon Agents
paperYour notes
Alibaba Token Hub's RL infrastructure for what the paper calls xLong-horizon agent training, where "a single execution can span hours, hundreds of model–environment interactions, and nearly 1M tokens per rollout." Two properties of that regime break ordinary RL systems. Rollout durations vary enormously, so fixed rollout/training GPU partitions idle (on NL2RepoBench with Qwen3.6 122B, mean rollout execution runs 1.93 hours and a query 2.96 hours; with Qwen3.8 2.4T, 61.1% of queries take at least four hours). And harness executions branch — through compaction, delegation and retries — so one execution yields many overlapping trajectories with duplicated prefixes. QwenGyre answers with an elastic scheduler that reallocates GPUs between rollout and training at cell granularity without interrupting live executions (requests reroute, KV cache migrates; switching costs 8.52 s into training and 3.46 s back, negligible against hour-long rollouts) and a trajectory processor that rebuilds each execution as a prefix-sharing trajectory tree, scores preserved work after a timeout, counts shared targets once, and caps trajectories per execution to bound training cost.
Experiments use an adaptation of GSPO with token-level importance weighting and execution-level clipping, 16 rollouts per group, and Claude Code 2.1.220 as the harness executing shell commands in containers. Against two baselines at matched GPU budgets and matched scheduling staleness — Async (fixed pools) and Colocate (one pool alternating phases) — QwenGyre reports end-to-end speedups of 1.38–1.78× over Async and 1.21–1.85× over Colocate while matching their training-score curves (DeepSWE 1.569–1.572 vs Async and 1.610–1.824 vs Colocate; TerminalBench 1.383–1.430 and 1.809–1.849). On the flagship Qwen3.8 2.4T at 700K tokens per rollout it completes 48 NL2RepoBench training steps in 75.42 hours against 134.55 for Async and 91.47 for Colocate, lifting the evaluation pass rate from 52.5% to 58.5%. Twelve authors led by Alibaba Token Hub with USTC and Tsinghua co-affiliations; corresponding author JianWei Zhang. No code release accompanied the paper.