arXiv AI By Tianyu Chen, Mingyuan Zhou, Jiaxing Wu

Learning to Route in Visual Space via Multi-Step Embedding Retrieval

Read the original on arXiv AI →

The paper introduces VHOP, a data generation framework and benchmark for visual agentic search, and VHOP-Router, an end‑to‑end training pipeline that turns a standard embedding model into an autoregressive multi‑step retriever operating directly in visual latent space. VHOP-Router eliminates the need for intermediate text queries, boosting retrieval accuracy from under 5% to 76.3% and improving task success rates by 52.7% while dramatically reducing token usage and API payloads. The approach generalizes to unseen difficulty levels and realistic test sets, offering an efficient solution that preserves native LLM capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 1

Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering

arXiv:2604.07146v3 Announce Type: replace Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for...

By Zhuohong Chen, Zhenxian Wu, Yunyao Yu, Hangrui Xu, Zirui Liao, Zhifang Liu, Xiangwen Deng, Pen Jiao, Haoqian Wang
arXiv AI
1d ago

EviRover: Reinforcing Agentic Perception Beyond a Glance

EviRover is a perception agent that goes beyond a single glance by actively gathering information to resolve perceptual queries. The authors created two data generation pipelines, producing EviRover-SFT-5K and EviRover-RL-12K, and a human‑verified benchmark called EviLens with 688 instances across five perception categories. Trained with supervised fine‑tuning and agentic reinforcement learning, the 4B EviRover outperforms its backbone by an average of 30 points on EviLens and shows strong transfer to other benchmarks such as WebEyes and BrowseComp‑VL.

By Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
arXiv AI
Aug 5

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

arXiv:2608. 03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration.

By Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao
arXiv AI
Sep 1

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

WeAgent-MMSearch introduces a multimodal search agent that preserves retrieved images as persistent references, enabling the model to inspect, process, and cite them throughout a search trajectory. The system includes a harness (WeAgent-Harness), a post‑training method (FA‑GSPO) that recovers salvageable rollouts, and a new benchmark (VisTarget‑Bench) to evaluate image‑retrieval versus visual‑perception failures. Evaluation shows that agentic post‑training boosts performance by 19.22 points, allowing the model to outperform similarly sized open‑source models and compete with much larger ones.

By Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao, Han Li, Juntao Liu, Wenchao Chen, Chengduo Zhao, Chao Yu, Fandong Meng