arXiv AI By Yuning Han, Yangchenchen Jin, Tyler Jandreau, Jingwei Sun

LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents

Read the original on arXiv AI →

LatentSift is a token‑free, execution‑free filtering method for software engineering agents that replaces the initial LLM‑based verifier. It represents each candidate trajectory using the policy’s hidden states—reasoning, observation, and function‑call states—and compares them against banks of successful and unsuccessful states collected during training. By fusing distance scores with a learned linear score, LatentSift retains promising candidates, reducing EF‑verifier tokens by 66.6–81.0% and total verification tokens by 49.1–62.1% while maintaining or improving performance on DeepSWE‑Preview.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

arXiv:2606. 07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.

By Rishi Desai, Jesse Hu, Joan Cabezas, Neel Harsola, Pratyush Shukla, Roey Ben Chaim, Adnan El Assadi, Omkaar Mukund Kamath, Fenil Faldu, Prannay Hebbar, Jiankai Sun, Yiyuan Li, Pramod Srinivasan, Ishan Gupta, Christopher Settles, Daniel Wang, Derek Chen, Pranav Raja, Albert Liu, Marek \v{S}uppa, Nevasini Sasikumar, Luyang Kong, Erik Quintanilla, Xiangyi Li, Ivan Bercovich, Steven Dillmann
arXiv Machine Learning
Jul 31

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

arXiv:2607. 28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks.

By Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai