arXiv AI

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

arXiv:2608. 11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway.

arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
2d ago

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

The paper introduces Rules to Tools (R2T), a system that provides executable checks for scientific coding agents to verify compliance with public scientific requirements. In experiments across multiple task cohorts, agents using R2T’s prepared checks achieved high repair success rates—26/30 with text and 29/30 with checks—while also demonstrating varying task preferences and cost trade‑offs. The study quantifies how tool‑enabled checks influence repair outcomes and agent‑side resource usage in scientific computing contexts.

By Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke
arXiv AI
Aug 28

Same Model, Different Harness: Different Coding-Agent Results

The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.

By Sydney Lewis
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran