Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments. Reliable estimates would therefore allow environment designers to calibrate evaluation benchmarks and construct progressive training curricula.
arXiv:2604. 00594v2 Announce Type: replace Abstract: As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficult.
By Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, Kaivalya Hariharan
The paper investigates what makes software issue resolution tasks difficult for agents by proposing a measurement framework and conducting a large‑scale empirical study on the CoderForge‑Preview dataset. It extracts static features from task patches, repositories, and prompts, and uses ensemble methods, SHAP attribution, and effect size analysis to predict task outcomes. The study finds that task difficulty is largely predictable from static features (AU C = 0.863), driven mainly by patch fragmentation and repository scale, with prompt linguistic features contributing for mid‑band tasks, suggesting a layered difficulty structure.
By Ebtesam Al-Haque, Brittany Johnson
arXiv:2602. 07267v2 Announce Type: replace Abstract: Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty.
By Fengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau, Siva Reddy, Hugo Larochelle
EarlyEval introduces a lightweight framework that predicts an LLM agent’s final outcome early in its execution, allowing the run to halt when a LightGBM classifier reaches a calibrated confidence threshold. By training success and failure classifiers on behavioral, textual, and reference-solution features, EarlyEval can cut 13%-26% of agent steps and up to 44.1% of input tokens while maintaining 89%-97% prediction accuracy. Across three benchmarks—SWE-bench Verified, TerminalBench, and Toolathlon—this approach reduces evaluation costs with minimal impact on per-agent resolve rates.
By Yuling Shi, Zhensu Sun, Junsen Dong, Chengcheng Wan, David Lo, Xiaodong Gu
arXiv:2608. 08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts.
By Dongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan