UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
arXiv:2607. 26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks.
arXiv:2608. 03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet.
arXiv:2607. 26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks.
The paper introduces ANASSA, an agentic AI orchestration framework designed for spatial intelligence in geographic information systems. It addresses gaps in current systems by integrating structured spatial reasoning, multi‑agent workflow orchestration, execution feedback, authoritative validation, provenance, uncertainty handling, and human decision authority. The architecture is detailed with eleven components across four layers, a six‑step Geospatial AI Cognitive Loop, cross‑component contracts, and governance mechanisms to ensure traceability, reproducibility, and accountability.
arXiv:2608. 08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.
arXiv:2606. 11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems.
Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.
arXiv:2606. 10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge.
arXiv:2606. 01725v1 Announce Type: new Abstract: Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes.
arXiv:2607. 13558v1 Announce Type: new Abstract: Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic assessment, and environmental monitoring.
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation.
The paper introduces GMA, a new benchmark for evaluating general mobile assistants in realistic, challenging scenarios. GMA expands on existing benchmarks by offering seven open‑source applications across diverse domains and 300 tasks organized into four difficulty tiers, ranging from simple actions to complex multi‑step workflows. The authors evaluate eight state‑of‑the‑art models, showing that performance drops sharply with task complexity, and conduct ablation studies on harness design—such as context retention and state tracking—to demonstrate how these choices can improve outcomes, especially for demanding workflows.
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.