arXiv Computation and Language

Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

arXiv AI
Sep 3

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

UniToolCall introduces a unified framework for tool-use in large language model agents, standardizing toolset construction, dataset generation, and evaluation. The framework aggregates over 22,000 tools and creates a hybrid training corpus of more than 390,000 instances by combining ten public datasets with synthetically generated, structurally controlled trajectories. It models diverse interaction patterns—single‑hop vs. multi‑hop, single‑turn vs. multi‑turn, serial vs. parallel execution—and adds an Anchor Linkage mechanism to enforce cross‑turn dependencies, while converting seven public benchmarks into a common Query–Action–Observation–Answer format for fine‑grained evaluation.

By Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
arXiv AI
Aug 28

SWE-Prime: Fewer Trajectories, Better Performance

SWE-Prime introduces a two-stage supervised fine-tuning data selection process for large language models tackling software issues. The first stage filters entire trajectories by quality and representativeness, while the second stage selects meaningful semantic segments based on contribution, learnability, and risk. Experiments on SWE-Bench Pro and Verified demonstrate that training on just 10% of trajectories chosen by SWE-Prime surpasses full-dataset training, achieving up to 12.2% and 24.2% performance gains.

By Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng
Hugging Face Trending Papers
Aug 27

SWE-Prime: Fewer Trajectories, Better Performance

SWE-Prime introduces a two‑stage, multi‑granularity supervised fine‑tuning (SFT) data selection process for large language models tackling real‑world software problems. The first stage filters entire trajectories based on process quality, result quality, and representativeness, while the second stage evaluates semantic segments for contribution, learnability, and risk, keeping all segments in context but only penalizing selected ones during training. Experiments on SWE‑Bench Pro and Verified demonstrate that training on just 10% of trajectories chosen by SWE‑Prime outperforms full‑dataset training, achieving up to 12.2% and 24.2% relative gains.

arXiv AI
Aug 26

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Selective Regenerative Decoding (SRD) is a new inference-time decoding method that improves large language model reasoning by allowing segment-level intervention on candidate trajectories. Instead of discarding or keeping entire trajectories, SRD selectively refines only the degraded suffix while preserving useful prefixes, leading to higher expected trajectory quality and better sample efficiency. Experiments on MATH500, GPQA Diamond, HotpotQA, and AlpacaEval show that SRD matches Best-of-N accuracy with fewer generated tokens and outperforms speculative rejection in low‑compute settings.

By Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung, Yi-an Lai, Arshit Gupta