arXiv:2608. 03609v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data.
By Alejandro J. Mercado, Alessio Lomuscio
Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.
By Harish Karumuri, Mahesh Vemula, David Lopes Pegna
arXiv:2606. 20615v3 Announce Type: replace Abstract: AI agents now act as first-class members of the software development lifecycle, but the instruments teams use to direct them enforce nothing: process encoded in prompts is flexible but unenforceable, while workflow formalisms are enforceable but do not model autonomous agents.
By Ylli Prifti, Pasquale De Meo, Alessandro Provetti
arXiv:2606. 14790v1 Announce Type: cross Abstract: LLM-based multi-agent systems increasingly coordinate planning, reasoning, tool use, and human interaction, yet their reliability remains limited.
By Hanqi Li, Jing Peng, Zijian Wang, Lu Chen, Kai Yu
arXiv:2606. 06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence.
By Ruida Wang, Jerry Huang, Pengcheng Wang, Xuanqing Liu, Luyang Kong, Tong Zhang
arXiv:2605. 10555v2 Announce Type: replace Abstract: As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-oriented CRUD paradigms.
By Kai Pan, Rong Hou
arXiv:2607. 16266v1 Announce Type: cross Abstract: Existing approaches for reasoning about action and change provide expressive semantics for modeling dynamic systems, in most cases built on top of logic programming systems.
By Julian Alfredo Mendez, Andreas Br\"annstr\"om
arXiv:2607. 14456v1 Announce Type: cross Abstract: Large Language Models (LLMs) have accelerated the adoption of software development agents, now widely available as Integrated Development Environment (IDE) extensions and standalone applications.
By Harris Borman, Herman Wandabwa, Fusun Yu, Sandeepa Kannangara, Justin Liu, Anna Leontjeva, Ritchie Ng
The paper proposes a four‑dimensional formal framework—Semantic Expressivity, Agentic Discoverability, Task‑Relative Grounding, and Epistemic Trust Scope—to extend current KG metadata standards (VoID and DCAT). It introduces the Agentic Affordance Profile (AAP), a semantic layer that enables agents to select, compose, and diagnose failures in knowledge graphs at planning time. A scholarly‑search example illustrates the framework and outlines a five‑point research agenda for scaling AAP‑based affordance matching.
By Terry R. Payne, Valentina Tamma, Enrico Daga
arXiv:2605. 09045v2 Announce Type: replace Abstract: Agentic frameworks are the software layer through which AI agents act in the world.
By Royce Moon, Lav R. Varshney
arXiv:2604. 22455v2 Announce Type: replace Abstract: A core component of any AI-Augmented Business Process Management System (ABPMS) is the process frame, which gives the system process-awareness and defines its maximal behavioral boundaries.
By Anti Alman, Izack Cohen, Avigdor Gal, Fabrizio Maria Maggi, Marco Montali
arXiv:2606. 05043v1 Announce Type: new Abstract: The last few years have witnessed major advances in the modeling and implementation of multiagent systems based on declarative interaction protocols.
By Samuel H. Christie V, Amit K. Chopra, Munindar P. Singh