What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The article reviews how replacing components in language‑model agents affects execution trajectories and downstream outcomes. It maps 348 studies, analyzes 90 comparison records, and finds that while many studies report both local decision metrics and task endpoints, they rarely demonstrate matched comparisons or prove that improved local decisions drive task‑level gains. The review identifies three potential mechanisms—recovery and disruption, intervention timing, and downstream use—and proposes eight claim‑specific reporting items to clarify evidence quality.
The paper introduces NAQD‑Env, a synthetic benchmark designed to test language agents’ ability to selectively withdraw and resume actions when new evidence, permissions, or stop instructions arise. It evaluates models against a deterministic reference policy across eleven dependency families, measuring policy agreement, task value, withdrawal, resumption, and event reporting. Experiments on 350 scenarios show low withdrawal recall, no valid resumption, and limited policy alignment, highlighting the need to treat selective withdrawal as a distinct reliability component.
The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.
The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.
arXiv:2609.36138v1 Announce Type: new Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or...
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.