arXiv:2603.19896v2 Announce Type: replace
Abstract: Tool-using large language model (LLM) agents often face a fundamental tension between answer quality and execution cost. Fixed workflows are stable...
By Boyan Liu, Gongming Zhao, Hongli Xu
arXiv:2610.02330v1 Announce Type: new
Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condi...
By Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng
arXiv:2609.06835v1 Announce Type: cross
Abstract: Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these s...
By Chaoyu Zhang, Hexuan Yu, Heng Jin, Shanghao Shi, Ning Zhang, Yi Shi, Yulia R. Gel, Y. Thomas Hou, Wenjing Lou
arXiv:2610.11899v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-a...
By Irene Weber (University of Applied Sciences Kempten, Germany)
arXiv:2606. 30555v1 Announce Type: new Abstract: The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows.
By Dvir Alsheich, Adar Peleg, Ben Hagag, Rom Himelstein, Amit Levi, Avi Mendelson
arXiv:2610.08155v1 Announce Type: cross
Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by en...
By Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang
arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.
By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
OrchSLM is a routing framework that unifies non‑interactive orchestration methods for small language models (SLMs). It allows heterogeneous SLMs to independently generate candidate solutions while a router manages their cached outputs without further model interaction. By systematically probing OrchSLM, the study shows how orchestration behavior depends on task structure, model‑pool composition, and multi‑agent consensus.
By Chengxi Zhang, Yu Yao
arXiv:2607. 03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.
By Alexander Somma, Isabelle Plante, Fred Premji
Sapien is a policy engine that enforces stateful contextual policies for autonomous AI agents, specifying allowed tool‑call sequences with an extended regular expression that includes stateful predicates, deferred policy generation, and scoped semantic checks. The system maintains performance close to an unconstrained agent while significantly reducing malicious actions, ruling out 93‑95% of attacks on AgentDojo and 62‑85% on Toolathlon, outperforming traditional tool allowlists on long‑horizon tasks.
By Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai
arXiv:2606. 09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite.
By Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran