arXiv AI

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan
arXiv AI
Sep 12

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

The study examines the composition of a random sample from the Model Context Protocol (MCP) registry, revealing that only 48.8% of the 400 sampled npm/stdio servers successfully complete an initialization handshake, compared to 66.7% for a hand‑curated frame. Among the servers that run, there are no fatal JSON Schema violations across 2,766 advertised tools, but optional safety annotations vary widely, with a 58.8% omission rate in the random draw versus 41.5% in the curated set. The authors also compare MCP tool descriptions to two benchmark corpora, finding minimal near‑duplication in real MCP tools (2.8%) and significant repetition in synthetic datasets (up to 85.6%).

By Haseeb Mohammed Afsar
arXiv Computation and Language
Aug 21

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

arXiv:2608. 19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling.

By Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Olafsson, Tommy Guy
arXiv AI
Jul 24

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya