The paper introduces VHOP, a data generation framework and benchmark for visual agentic search, and VHOP-Router, an end‑to‑end training pipeline that turns a standard embedding model into an autoregressive multi‑step retriever operating directly in visual latent space. VHOP-Router eliminates the need for intermediate text queries, boosting retrieval accuracy from under 5% to 76.3% and improving task success rates by 52.7% while dramatically reducing token usage and API payloads. The approach generalizes to unseen difficulty levels and realistic test sets, offering an efficient solution that preserves native LLM capabilities.
By Tianyu Chen, Mingyuan Zhou, Jiaxing Wu
arXiv:2603.18272v2 Announce Type: replace
Abstract: While large language models (LLMs) have advanced the development of general-purpose agents, robust generalization to unseen tasks remains challengi...
By Thomas Palmeira Ferraz, Romain Deffayet, Vassilina Nikoulina, Herv\'e D\'ejean, St\'ephane Clinchant
TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.
By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan
arXiv:2608. 12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed.
By Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu
arXiv:2604. 23336v3 Announce Type: replace-cross Abstract: Unlike traditional fact-based retrieval, rationale-based retrieval typically necessitates cross-encoding of query-document pairs using large language models, incurring substantial computational costs.
By Teng Chen, Sheng Xu, Feixiang Guo, Xiaoyu Wang, Qingqing Gu, Hongyan Li, Luo Ji
arXiv:2601.21545v2 Announce Type: replace
Abstract: Agentic systems accumulate persistent memory across sessions, tools, and tasks, and a later request must retrieve from it under two distinct constr...
By Yang Zhao, Chengxiao Dai, Mengying Kou, Yue Xiu, Dusit Niyato