arXiv Machine Learning By Hamidah Oderinwale

Agent trajectories as programs: fingerprinting and programming coding-agent behavior

Read the original on arXiv Machine Learning →

arXiv:2606. 16988v1 Announce Type: cross Abstract: Benchmark scores tell you what an agent got right; they do not tell you how it got there.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf is a new semantic profiler designed for long‑horizon AI agents that aggregates agent trajectories into pprof‑compatible profiles, enabling flame‑graph visualization and hierarchical attribution of tasks and subtasks. It introduces a semantic operation stack model and recursive operation segmentation to replace traditional call‑stack profiling, addressing the challenge of profiling agent intent rather than code paths. In evaluations, AgentPProf achieves high F1 scores against human annotations and significantly improves problem‑localization metrics, demonstrating its effectiveness in attributing resources, locating issues, and optimizing token cost.

By Yusheng Zheng, Chaokun Chang, Yu Mao, Tianyuan Wu, Yuxi Huang, Tao Ma, Wenan Mao, Shuyi Cheng, Andi Quinn, Wei Wang
arXiv AI
Aug 20

What Makes Software Issue Resolution Tasks Difficult for Agents?

The paper investigates what makes software issue resolution tasks difficult for agents by proposing a measurement framework and conducting a large‑scale empirical study on the CoderForge‑Preview dataset. It extracts static features from task patches, repositories, and prompts, and uses ensemble methods, SHAP attribution, and effect size analysis to predict task outcomes. The study finds that task difficulty is largely predictable from static features (AU C = 0.863), driven mainly by patch fragmentation and repository scale, with prompt linguistic features contributing for mid‑band tasks, suggesting a layered difficulty structure.

By Ebtesam Al-Haque, Brittany Johnson