arXiv AI

Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)

arXiv:2606. 05145v1 Announce Type: cross Abstract: When post-trained language models fail on reasoning problems, the common test-time-scaling response is to spend more compute on additional attempts, and the failed traces play no further role.

arXiv AI
Aug 11

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents

arXiv:2608. 08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts.

By Dongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan
arXiv AI
3d ago

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

The Agent Error Dataset (AED) presents 50,228 error–diagnosis pairs collected from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text‑based agent systems. A five‑stage Agentic Error‑to‑Training (AET) pipeline generates diagnoses and proposed corrections, verifies them against recorded evidence, and creates separate training views for diagnosis and actor recovery. Experiments show that first‑proposal corrections improve verifier pass rates from 18.4% to 51.1%, and fine‑tuning with full‑diagnosis data raises Qwen3‑8B’s exact‑step agreement from 47.2% to 63.6% on a holdout set.

By Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
arXiv AI
Aug 13

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

arXiv:2608. 11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification.

By Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Junlin Zhang
arXiv AI
Jun 9

REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

arXiv:2606. 09071v1 Announce Type: new Abstract: Large language model (LLM) agents now solve complex tasks through long plan-and-execution traces, yet the ability to locate errors in a completed traces still lags far behind, especially in the \emph{silent failure} regime.

By Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok, Daniel Guo, Sahil Arun Nale, Charles Fleming, Guang Cheng
arXiv AI
Sep 2

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

The study investigates whether hints that convert failing code generation attempts into passing ones provide new information or simply guide models toward solutions they could already generate. Using Qwen2.5-3B-Instruct and Phi-3.5-mini on HumanEval+ and MBPP+, the authors find that relevant hints rescue a significant portion of failures, yet many of those solutions are also recoverable through ordinary sampling. Mechanistic tests reveal a shared activation direction between relevant and unrelated hints, but adding this direction does not improve overall accuracy, indicating limited task-general transfer.

By Will Badr