HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
arXiv:2608. 02009v2 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence.
arXiv:2608. 02009v2 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence.
arXiv:2609.15982v1 Announce Type: cross Abstract: Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by pre...
Grounded Continuation introduces a runtime verifier that classifies each utterance in an LLM conversation into one of eight epistemic operations and uses a symbolic engine to maintain a dependency map of claims and their supports. The verifier checks whether a new continuation is grounded by walking this map, a linear-time process that requires no additional LLM calls. On benchmarks such as ReviseQA and MemoryAgentBench, the verifier improves single-hop accuracy for several QA models, even enabling a 7B model to outperform GPT‑4o when guided by the verifier.
The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.
The paper introduces LAST-CQ, a five-agent, training‑free, execution‑grounded framework for Text‑to‑Cypher that evaluates which components of an agentic pipeline contribute most to performance. Experiments on 2,471 live‑database queries across six backbones show that removing correction reduces execution‑BLEU by 3.1–12.3%, while substituting schema‑grounded feedback with raw error strings has negligible impact. Parallel sampling degrades quality by 10–11%, whereas failure detection and retry routing recover 91.7% of initially failed queries, highlighting that simple failure handling is more effective than sophisticated feedback or increased sampling.
arXiv:2608. 04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone.
arXiv:2609.22223v1 Announce Type: cross Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing c...
The paper presents a rigor‑matched audit comparing two periodic‑step, search‑based layer‑skipping methods for efficient large language model inference: a confidence‑gated early‑exit baseline (ConfLayers) and a self‑speculative decoding approach (SWIFT). Across two Qwen2.5 model scales and tasks (GSM8K reasoning and CNN/DailyMail summarization), SWIFT consistently outperforms ConfLayers in accuracy and, after separating search overhead, achieves faster true inference speed in most settings. The study also evaluates two trained‑routing methods (LayerRoute and LayerDrop), finding modest speedups but significantly lower accuracy, especially for LayerRoute on GSM8K at 1.5B.
arXiv:2608. 16391v1 Announce Type: cross Abstract: As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem.
arXiv:2606. 08151v1 Announce Type: new Abstract: Tool-using LLM agents often fail not because relevant text is absent, but because decisive evidence is not selected, compressed, or surfaced at action time.
The paper introduces Traverse, a benchmark of 2,518 agent trajectories and 6,967 annotated mistakes across software engineering, computer use, and science tasks, revealing that failures often go unrecovered and can cause irreversible harm before a run is deemed successful. It shows that human judges struggle to detect the first mistake in most runs, while a 4‑billion‑parameter verifier called Scout can locate failures more effectively and improve task success when used to select among candidate runs. The study demonstrates that making failure detection inexpensive and reliable can enable long‑horizon agents to learn from their own mistakes and increase trustworthiness in autonomous AI.
arXiv:2607. 12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy.