arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan
arXiv AI
Sep 12

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

The study examines the composition of a random sample from the Model Context Protocol (MCP) registry, revealing that only 48.8% of the 400 sampled npm/stdio servers successfully complete an initialization handshake, compared to 66.7% for a hand‑curated frame. Among the servers that run, there are no fatal JSON Schema violations across 2,766 advertised tools, but optional safety annotations vary widely, with a 58.8% omission rate in the random draw versus 41.5% in the curated set. The authors also compare MCP tool descriptions to two benchmark corpora, finding minimal near‑duplication in real MCP tools (2.8%) and significant repetition in synthetic datasets (up to 85.6%).

By Haseeb Mohammed Afsar