arXiv AI

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

The paper introduces RUPA, a trajectory‑level uncertainty quantification framework for large language model agents. RUPA models an agent’s execution as a directed graph of reasoning states, tool interactions, and environment feedback, then propagates uncertainty across this graph to capture long‑range dependencies. Experiments on benchmarks such as τ‑2, Terminal‑Bench‑2, and GAIA show that RUPA outperforms existing methods, enabling earlier failure detection and more reliable agent execution.

arXiv Computation and Language
Aug 25

PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems

PropUQ-MAS is a framework for uncertainty quantification in large language model (LLM) multi‑agent systems that models the system as a communication‑structured graph. It estimates the reliability of each step by combining local uncertainty with uncertainty inherited from upstream messages, addressing the risk of error propagation in inter‑agent communication. Experiments show consistent improvements in UQ metrics, with average gains of +6.10% in AUROC and +47.58% in PRR.

By Yaokun Liu, Yifan Liu, Daniel Yue Zhang, Ruichen Yao, Zelin Li, Dong Wang
arXiv AI
Sep 4

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It contains 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency checking, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.

By Yaxing Lyu, Shengjie Zhou, Binbin Toh, Pengyu Zhu, Lijun Li
arXiv Machine Learning
Sep 24

Learning from Failures: Heterogeneous Graph Memory for Small Language Model Tool-Using Agents

The paper introduces FRESH, a Failure-aware Retrieval framework that uses Experience-Structured Heterogeneous graphs to transform past successes and failures into structured external memory for tool‑using agents. By explicitly modeling dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH enables frozen language models to reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on τ‑Bench and AppWorld with multiple open‑source models demonstrate that FRESH consistently improves task success and tool‑use reliability compared to no‑memory agents and other memory‑based baselines.

By Jiaxing Li, Lei Song, Rui Dong, Youyong Kong
Hugging Face Trending Papers
Sep 3

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It consists of 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.