arXiv AI

SheetMind: Actions Set Accuracy, Agents Set the Failure Mode

SheetMind is a Manager‑Action‑Reflection framework that evaluates how much spreadsheet agent performance derives from the agents themselves versus the shared action interface. In a controlled study on all 221 tasks of the SheetCopilot Benchmark, replacing the high‑level action API with primitive cell operations drops accuracy by 47.1 points, while adding a Reflection Agent improves performance by 4.5 points and a Manager by 1.4 points. The framework also shows that decomposition changes failure modes, reducing silent wrong outputs from 33% to 25%, and that GPT‑5 and GPT‑5‑mini achieve similar performance, whereas GPT‑3.5 underperforms significantly.

arXiv AI
Sep 15

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

The paper introduces a new evaluation method called "same-input rerun" to assess the consistency of clinical language‑model agents across repeated runs. By replaying 1,000 MedAgentBench tasks with identical inputs, the authors find that action‑level outputs—such as test orders, medication requests, and referrals—vary significantly, even when benchmark scores remain unchanged. The study demonstrates that current benchmarks, which typically evaluate only a single run per task, can miss substantial behavioral divergence.

By Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang
arXiv AI
Aug 25

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.

By Pedro Santos
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Sep 25

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.

By Ivan Matveev
arXiv AI
Sep 11

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

AgentAudit is an open, extensible framework that evaluates the full lifecycle of AI agents, assessing planning, tool selection, execution, memory, and reasoning across ten dimensions such as instruction integrity, security, and alignment. Unlike existing benchmarks that focus on single aspects, AgentAudit analyzes the entire execution trace to attribute failures to specific stages. The framework was tested on five large language models, revealing significant differences in trustworthiness even among models with similar task‑completion performance.

By Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar