arXiv AI By Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, Kaivalya Hariharan

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Read the original on arXiv AI →

arXiv:2604. 00594v2 Announce Type: replace Abstract: As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficult.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

What Makes Software Issue Resolution Tasks Difficult for Agents?

The paper investigates what makes software issue resolution tasks difficult for agents by proposing a measurement framework and conducting a large‑scale empirical study on the CoderForge‑Preview dataset. It extracts static features from task patches, repositories, and prompts, and uses ensemble methods, SHAP attribution, and effect size analysis to predict task outcomes. The study finds that task difficulty is largely predictable from static features (AU C = 0.863), driven mainly by patch fragmentation and repository scale, with prompt linguistic features contributing for mid‑band tasks, suggesting a layered difficulty structure.

By Ebtesam Al-Haque, Brittany Johnson
Hugging Face Trending Papers
Aug 6

Predicting Task Difficulty Without Rollouts

Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments. Reliable estimates would therefore allow environment designers to calibrate evaluation benchmarks and construct progressive training curricula.

arXiv Machine Learning
Aug 7

Predicting Task Difficulty Without Rollouts

arXiv:2608. 05797v1 Announce Type: new Abstract: Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments.

By Stefan Krsteski, Charlotte Meyer
arXiv AI
Sep 4

Efficient Test-Time Adaptation through Human-AI Interaction

The paper introduces Test-Time Adaptation through Human‑Agent Interaction (TAHI), a method that uses iterative human feedback to adapt AI agents to individual users’ criteria. By integrating cross‑session interaction data into agent context and weights, and building an evolving rubric module, the authors demonstrate that agents can improve task success by 4.5–20.9% after only a few interactions. The evolving rubric also serves as a scalable annotation tool, detecting 16.0–22.3% more failures than language models or humans alone, and personalized agents can even generalize improvements up to 8.8% across users.

By Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, Daniel Fried