arXiv AI

LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation

arXiv:2602. 16953v3 Announce Type: replace Abstract: Execution-aware LLM agents offer a promising paradigm for learning from tool feedback, but such feedback can be expensive and slow to obtain, making online reinforcement learning (RL) less practical in certain scenarios.

arXiv Computation and Language
Sep 18

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

CovR is an agentic framework that automates testbench generation for hardware verification by combining self-reflection loops with simulation-based feedback to maximize coverage. It builds a large dataset of 16,514 specification–RTL reasoning tuples and uses reinforcement learning with tool-derived rewards to train a student model, achieving high coverage scores on VerilogEval, RTLLM V2.0, and CVDP. When deployed as a plug-in stimulus engine, CovR boosts coverage by nearly 19% and improves mutation detection while uncovering previously undetected failures.

By Manar Abdelatty, Maryam Nouh, Sherief Reda
arXiv AI
Jul 7

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

arXiv:2602. 21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks.

By Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang
arXiv AI
Jul 20

ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

arXiv:2607. 15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration.

By Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng
arXiv AI
Aug 18

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.

By Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
arXiv AI
Jun 3

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

arXiv:2606. 03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to build, synthetic training queries are often detached from the server's actual state (so the generated tool calls fail to execute), and recall-based RL rewards incentivize verbose tool-calling patterns.

By Ibrahim Abdelaziz, Asim Munawar, Kinjal Basu, Maxwell Crouse, Chulaka Gunasekara, Suneet Katrekar, Pavan Kapanipathi
arXiv Machine Learning
Aug 27

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

TailSFT is a simple modification to supervised fine‑tuning that filters out already well‑modeled sequences, concentrating learning on the tail of the data distribution. On the OLMo‑3 7B model, this approach improves pass@16 performance on math and coding tasks by up to 17% absolute and yields up to 4% absolute gains in subsequent GRPO reinforcement‑learning runs, with only minimal computational overhead. The authors also provide a lightweight diagnostic to identify settings where TailSFT is most beneficial and argue for a stage‑aware development strategy that evaluates intermediate checkpoints by their support for later training.

By Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy