Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficiently understood. This paper reports a reproducible phenomenon observed in a production Agent system: when Tool Calling and JSON Schema constraints are simultaneously enabled, multiple open-weight models cease invoking tools despite maintaining high schema compliance.
arXiv:2607. 00035v1 Announce Type: new Abstract: LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures.
By Bo Chen
arXiv:2609.23742v1 Announce Type: new
Abstract: Small open-source large language models (LLMs) in the 0.6B-4B parameter range are increasingly deployed for structured output generation (JSON, functio...
By Akash Chavan
arXiv:2605.23916v2 Announce Type: replace-cross
Abstract: AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descrip...
By Haochuan Kevin Wang, Zechen Zhang
We are introducing Structured Outputs in the API—model outputs now reliably adhere to developer-supplied JSON Schemas.
arXiv:2605.29313v2 Announce Type: replace
Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult...
By Shuyu Zhang, Yaqi Shi, Jiarui Zhang, Yanxiao Zhao, Lu Wang
arXiv:2608. 13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation.
By Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li
arXiv:2607. 29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions.
By Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan, Yu Jiang, Zhenpeng Chen
arXiv:2608.29128v1 Announce Type: new
Abstract: Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This metric fails to distinguish failures that matt...
By Zelin Wan, Arash Nourian, Xiaoxiao Li, Nihar Nandan, Kamalakannan Nandagopal
arXiv:2608. 14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools.
By Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert
The paper introduces a 202-scenario benchmark to evaluate how large language models (LLMs) handle safety-critical authorization decisions for vehicle voice commands. It tests two local open-weight models and three API-based LLMs, finding alignment scores ranging from 40.1% to 89.1% and noting persistent false execution errors. The study concludes that structured LLM decisions alone are insufficient for safety, recommending an independent enforcement layer to verify tool permissions and vehicle-state constraints before any vehicle function is invoked.
By Diba Afroze, Xingli Zhang, Yazhou Tu, Xiali Hei
arXiv:2609.26693v1 Announce Type: new
Abstract: A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action....
By Lijuan Tang, Yuemeng Zheng