arXiv AI

Environment-free Synthetic Data Generation for API-Calling Agents

arXiv:2607. 16900v1 Announce Type: new Abstract: Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories.

arXiv AI
Aug 20

Looped Language Models Improve Compositional Tool Calling

The paper investigates the use of looped language models for compositional tool calling, where models must coordinate multiple API calls and maintain state across interactions. Experiments on API-Bank, BFCL, and NESTful show that recurrent computation generally improves compositional and dependency-aware tool use, with accuracy increasing as recurrent depth grows. Adaptive inference offers a better compute‑performance trade‑off by allocating extra computation only when necessary.

By Andrei Cristian Popescu, Haitz S\'aez de Oc\'ariz Borde, Pietro Li\`o
arXiv AI
Jul 28

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

arXiv:2607. 23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes.

By Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
arXiv AI
Sep 7

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

The paper introduces KOPA-Bench, a benchmark of 145 real-world tasks that evaluate multi-step tool‑calling over Korean open public APIs. It also presents EDGE, a data‑synthesis method that builds an execution‑grounded dynamic graph to generate executable multi‑step trajectories, and shows that a fine‑tuned 9B model performs nearly as well as an untuned 27B model on KOPA‑Bench and the BFCL benchmark.

By Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim
arXiv AI
Jun 29

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.

By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
arXiv AI
Sep 1

SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning

SimCRAFT is a model‑agnostic framework that distills remote sensing orchestration into a compact 7B‑scale model. It creates a large, constraint‑validated workflow planning corpus (SimRS‑14k) using a multi‑agent synthesis engine and a Mock Execution Engine, then fine‑tunes the model with Contextual Retrieval‑Augmented Fine‑Tuning (CRAFT) to reason analogically. Experiments show SimCRAFT‑7B outperforms open‑weight LLMs and rivals advanced closed‑source models, providing a lightweight, efficient baseline for autonomous remote sensing deployment.

By Haoran Wang, Jing Yao, Xu Yang, Zeqing Wang, Yang Zhang, Pedram Ghamisi, Zhengchao Chen
arXiv AI
Sep 23

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

EDGEGEN is a synthetic task generation framework that extracts compliance rules from a tool‑calling agent’s specification to create database‑grounded edge‑case tasks that violate those rules. By combining EdgeGen with existing synthetic data generation methods, it forms a fully automated closed‑loop system that requires no human annotation. Experiments show that finetuning on EdgeGen data improves performance by 2–42 % on the tau2bench airline domain, while harness optimization yields 10–30 % gains over human‑curated and base harnesses for the Gemma‑4‑e4b model.

By Harshavardhan Abichandani, Penny Chong, Jiyuan Shen, Gunraj Singh, Ashutosh Hathidara, Marcus Duigan Xing Yu, Jane Lo, Atin Ghosh, Yipeng Li, Daniel Dahlmeier
arXiv AI
Jun 9

Exploring Autonomous Agentic Data Engineering for Model Specialization

arXiv:2605. 30407v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.

By Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng