arXiv AI

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

arXiv:2606. 21262v2 Announce Type: replace Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad.

arXiv AI
2d ago

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

RISED introduces a framework that uses rubric-based textual feedback to improve training of a single large language model (LLM) agent across multiple interactive environments. By having an LLM judge tag rollouts with a shared rubric vocabulary, the system guides both online data selection and policy supervision, enabling richer cross‑environment relationships and within‑group reward contrast. Experiments show that RISED achieves the highest mean pass rate and ranks first or second in every individual environment, with rubric analysis revealing behavioural changes behind these gains.

By Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu
arXiv AI
Sep 4

CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents

CoMAP introduces a framework that jointly evolves textual world models and agent policies through a closed‑loop interaction. At each decision step the world model forecasts future state feedback for candidate actions, while the agent reflects on the reliability of this feedback to refine its action. The resulting on‑policy trajectories are used to self‑distill and update the world model, improving prediction accuracy and long‑horizon decision‑making across embodied planning, web navigation, and tool‑use benchmarks.

By Youwei Liu, Jian Wang, Hanlin Wang, Wenjie Li
arXiv AI
Aug 26

Task-Adaptive Rubrics for GUI Reward Modeling

The paper introduces AdaptRubric, a Coarse-to-Fine Rubrics Framework designed to create task‑adaptive judging criteria for GUI reward modeling. It first retrieves a category‑level coarse rubric by mapping instructions to a GUI task family, then generates an instance‑level fine rubric that captures specific values, scopes, and constraints from the instruction. Experiments show that AdaptRubric outperforms existing reward agents, improving F1 by 3.6 points and achieving a 4.23‑point task‑success gain under a matched image budget.

By Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
arXiv AI
6d ago

CODESKILL: Learning Self-Evolving Skills for Coding Agents

CODESKILL is an LLM-based framework that learns to extract, evolve, and maintain procedural skills from coding-agent trajectories. It treats skill extraction and skill-bank management as a learnable policy trained with reinforcement learning, using a hybrid reward combining rubric-based skill quality and verifiable execution feedback. Experiments on EnvBench, SWE-Bench Verified, and Terminal-Bench 2 demonstrate that CODESKILL raises average pass rates by 11.03 over a no-skill baseline and by 5.10 over the strongest prompt-based or memory baseline while keeping a compact skill bank.

By Yanzhou Li, Yiran Zhang, Xiaoyu Zhang, Xiaoxia Liu, Yang Liu