arXiv AI

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.

arXiv Computation and Language
Aug 28

Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses

The paper investigates why large‑language‑model coding agents rarely request a second chunk of tool output, focusing on the precision‑at‑1 rate ($p_1$) of the gold item appearing first in the first chunk. In a benchmark of 500 software‑engineering tasks, the authors compare six value functions and find that increasing $p_1$ does not systematically improve downstream accuracy; the agent can recover the correct answer from any position within the chunk. Adding file‑metadata signals to a keyword scorer actually reduces $p_1$, while a parameter‑free keyword scorer improves $p_1$ but still fails to boost overall accuracy.

By Tatiana Petrova, Andrei Mazniak, Radu State
arXiv AI
Aug 28

Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.

By Alden Do Rosario, Hussein Younes, Felipe Pires
arXiv AI
3d ago

Can a Cacheable Decision Model Follow Rules?

The paper evaluates Certo, a small non‑generative decision model that scores candidate actions based on text. It compares a joint scorer that processes state, rules, and candidates together with a cacheable encoder that pre‑encodes candidates to reduce cost. Experiments show the cacheable approach loses rule sensitivity, while targeted counterfactual supervision can recover performance on synthetic tasks; however, on real rules the joint scorer still outperforms the cacheable version, and cross‑domain mixtures do not improve accuracy.

By Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP)
arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand