arXiv AI

Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System

The paper surveys how large language models (LLMs) are integrated into software systems under various labels such as chatbot, copilot, retrieval‑augmented generation (RAG), workflow, coding agent, and AI agent. It identifies seven recurring architectural forms—LLM chats, custom agents, RAG, AI‑enhanced workflows, copilots, coding agents, and agentic RAG—and describes each along four structural dimensions: architectural pattern, execution control, user intervention point, and tool usage. The study uses a corpus of 22 systems from research and vendor documentation to illustrate these forms and their limits.

arXiv AI
Sep 2

ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything

ChatDev 2.0, also called DevAll, is a no-code platform that lets users build, run, and inspect heterogeneous multi‑agent systems (MAS) powered by large language models. It combines a declarative executable graph abstraction with a cycle‑aware execution engine, enabling representation and execution of dynamic, cyclic interactions among diverse agents. The integrated visual interface allows users to author, monitor, and inspect MAS—including human‑in‑the‑loop steps—without writing code, and experiments show it matches state‑of‑the‑art MAS performance across three tasks.

By Yufan Dang, Shu Yao, Bowen Lai, Chenting Xu, Ruijie Shi, Wai-Shing Leung, Huatao Li, Chen Qian, Zhiyuan Liu
arXiv AI
Oct 2

Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems

The paper investigates challenges in open‑source large‑language‑model (LLM) based multi‑agent systems (MAS). By analyzing 944 issues extracted from 21 projects, it finds that orchestration and execution problems are most common, with workflow, tool integration, and memory issues as primary causes. The predominant remedy identified is optimizing workflow, and the study offers empirically grounded implications for improving orchestration, tool integration, and memory mechanisms in LLM‑based MAS.

By Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang, Zengyang Li, Arif Ali Khan
arXiv AI
Jun 15

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

arXiv:2606. 14502v1 Announce Type: new Abstract: Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-improvement.

By Yongheng Zhang, Ziang Liu, Jiaxuan Zhu, Shuai Wang, Xiangqi Chen, Haojing Huang, Jiayi Kuang, Siyu Chen, Ao Shen, Hao Wu, Qiufeng Wang, Qian-Wen Zhang, Junnan Dong, Wenhao Jiang, Ying Shen, Hai-Tao Zheng, Yinghui Li, Di Yin, Xing Sun, Philip S. Yu
arXiv AI
5d ago

Toward SLM-based agentic task-tool intent matching

The paper proposes using Small Language Models (SLMs) as a task‑tool relevance classifier to verify each tool call made by AI agents. By evaluating every selected tool against the assigned task, the SLM provides a relevance signal that can be used for downstream enforcement. The authors introduce a novel dataset of multi‑tool tasks across distinct Model Context Protocol servers and explore prompt‑optimization, supervised fine‑tuning, and reinforcement learning (GRPO) to optimize and specialize the SLMs.

By Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi
arXiv AI
Aug 18

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

arXiv:2608. 15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve.

By Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit