Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning¶
Source: converted from arXiv LaTeX source source/0.main.tex.
1.abstract
Figure fig:teaser: The performance of on reasoning and coding benchmarks. The four reasoning benchmarks are evaluated in a reasoning-with-Python-tool setting, where the baseline is the Qwen3-30B-A3B SFT model; SWE-Bench Verified evaluates coding with the Qwen3-30B-A3B baseline. outperforms the corresponding baseline and GRPO across all five benchmarks.
2.introduction 2.preliminary 4.method 5.experiment
3.related_work 6.conclusion
7.appendix