arXiv AI By Shuyan Huang, Kai Du, Andrew Lan

Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories

Read the original on arXiv AI →

arXiv:2608. 10319v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.

By Guanqun Yang, Wei Yang, Xueqing Liu