arXiv AI By Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Read the original on arXiv AI →

arXiv:2608. 11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
2d ago

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

The paper introduces Rules to Tools (R2T), a system that provides executable checks for scientific coding agents to verify compliance with public scientific requirements. In experiments across multiple task cohorts, agents using R2T’s prepared checks achieved high repair success rates—26/30 with text and 29/30 with checks—while also demonstrating varying task preferences and cost trade‑offs. The study quantifies how tool‑enabled checks influence repair outcomes and agent‑side resource usage in scientific computing contexts.

By Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke