arXiv AI

Toward SLM-based agentic task-tool intent matching

The paper proposes using Small Language Models (SLMs) as a task‑tool relevance classifier to verify each tool call made by AI agents. By evaluating every selected tool against the assigned task, the SLM provides a relevance signal that can be used for downstream enforcement. The authors introduce a novel dataset of multi‑tool tasks across distinct Model Context Protocol servers and explore prompt‑optimization, supervised fine‑tuning, and reinforcement learning (GRPO) to optimize and specialize the SLMs.

arXiv AI
Aug 10

An End-to-End Agent Auditing Engine

arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv AI
Sep 15

OrchSLM: Probing the Dynamics of Small Language Model Orchestration

OrchSLM is a routing framework that unifies non‑interactive orchestration methods for small language models (SLMs). It allows heterogeneous SLMs to independently generate candidate solutions while a router manages their cached outputs without further model interaction. By systematically probing OrchSLM, the study shows how orchestration behavior depends on task structure, model‑pool composition, and multi‑agent consensus.

By Chengxi Zhang, Yu Yao
arXiv AI
Oct 2

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Sapien is a policy engine that enforces stateful contextual policies for autonomous AI agents, specifying allowed tool‑call sequences with an extended regular expression that includes stateful predicates, deferred policy generation, and scoped semantic checks. The system maintains performance close to an unconstrained agent while significantly reducing malicious actions, ruling out 93‑95% of attacks on AgentDojo and 62‑85% on Toolathlon, outperforming traditional tool allowlists on long‑horizon tasks.

By Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai
arXiv AI
Jun 9

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

arXiv:2606. 09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite.

By Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran