arXiv:2610.01023v1 Announce Type: cross
Abstract: Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries...
By Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.
By Xiaonan Xu, Wenjing Wu
Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take, which failures are repaired, which states are verified, and which evidence is logged. We show that this harness can change the agent's multi-step beliefs even when the task, environment, and base LLM are fixed.
arXiv:2607. 04528v1 Announce Type: new Abstract: Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take, which failures are repaired, which states are verified, and which evidence is logged.
By Haiwen Yi, Xinyuan Song
arXiv:2609.35889v1 Announce Type: cross
Abstract: Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing...
By Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu
arXiv:2609.07095v1 Announce Type: new
Abstract: LLM assistants often produce more answers than humans can review before users see them. Most evaluations ask whether an answer is wrong, unsupported, o...
By SangJin Park, Myungsub Choi, Jineok Kim, Minseung Kang