arXiv AI

SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use

arXiv AI
Sep 3

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

UniToolCall introduces a unified framework for tool-use in large language model agents, standardizing toolset construction, dataset generation, and evaluation. The framework aggregates over 22,000 tools and creates a hybrid training corpus of more than 390,000 instances by combining ten public datasets with synthetically generated, structurally controlled trajectories. It models diverse interaction patterns—single‑hop vs. multi‑hop, single‑turn vs. multi‑turn, serial vs. parallel execution—and adds an Anchor Linkage mechanism to enforce cross‑turn dependencies, while converting seven public benchmarks into a common Query–Action–Observation–Answer format for fine‑grained evaluation.

By Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
arXiv Computation and Language
Aug 28

Agent Seer: Synthesizing Scenarios from Specification Understanding

Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.

By Harish Karumuri, Mahesh Vemula, David Lopes Pegna
arXiv AI
Sep 7

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

The paper introduces KOPA-Bench, a benchmark of 145 real-world tasks that evaluate multi-step tool‑calling over Korean open public APIs. It also presents EDGE, a data‑synthesis method that builds an execution‑grounded dynamic graph to generate executable multi‑step trajectories, and shows that a fine‑tuned 9B model performs nearly as well as an untuned 27B model on KOPA‑Bench and the BFCL benchmark.

By Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim
arXiv AI
Jul 7

Automated Data Readiness for Scientific AI

arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.

By Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi, Ketan Maheshwari, Marshall McDonnell, Massimiliano Lupo Pasini, Polina Shpilker, Renan Souza, Patrick Widener, Sarp Oral, Wesley Brewer
arXiv AI
Sep 3

Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

The paper introduces Tool Primitives, a design that replaces rigid API schemas with natural language interfaces for tool calling, enabling seamless inter-tool communication. It builds ToolFace, a repository of over 25,000 functions that LLMs can dynamically retrieve, and HEART, a harness engineering framework that orchestrates tool use with planning, routing, and verification. Experiments show HEART outperforms fine‑tuned models and leading commercial LLMs while cutting API costs by up to 85%.

By Haibo Jin, Suijin Wang, Xucheng Yu, Haojing Luo, Haohan Wang
arXiv AI
Aug 12

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

arXiv:2608. 11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use.

By Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantharaj, Sandro Fiore, Renan Souza