V‑Retrver is an evidence‑driven retrieval framework that treats universal multimodal retrieval as an agentic reasoning process grounded in visual inspection. It allows multimodal large language models to selectively acquire visual evidence through external tools, alternating between hypothesis generation and targeted visual verification. The approach is trained with a curriculum that blends supervised activation, rejection‑based refinement, and reinforcement learning, achieving an average 23.0% improvement in retrieval accuracy across multiple benchmarks.
By Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan
The paper introduces LARK, a two‑stage latent reasoning framework designed to mitigate cross‑modal dilution in multimodal recommendation systems. In the first stage, learnable latent tokens are interleaved with chain‑of‑thought reasoning and aligned with a frozen vision encoder to preserve visual details. The second stage projects these latent representations through a bridge MLP, employing item‑to‑item contrastive learning and aligning intermediate features with the first‑stage hidden states to anchor final embeddings to the model’s reasoning output. Experiments on three public benchmarks and an industrial dataset demonstrate that LARK achieves state‑of‑the‑art performance across multiple recommendation architectures, with ablation studies confirming the contribution of each component.
By Jiarui Jin, Anyang Ji
arXiv:2606. 15160v1 Announce Type: cross Abstract: Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years.
By David Huang, Lianlei Shan
arXiv:2610.01892v1 Announce Type: cross
Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...
By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
arXiv:2511. 17731v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs).
By Lingxiao Li, Yifan Wang, Xinyan Gao, Chen Tang, Xiangyu Yue, Chenyu You
ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.
By Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao
MMEmb-R1 is a multimodal embedding framework that enhances reasoning by treating it as a latent variable and selecting beneficial reasoning paths through pair-aware selection and counterfactual intervention. It uses reinforcement learning to invoke reasoning only when necessary, reducing unnecessary computation and latency. On the MMEB-V2 benchmark, MMEmb-R1 achieves a state‑of‑the‑art score of 71.2 with just 4 B parameters.
By Yuchi Wang, Dingkang Yang, Haiyang Yu, Weikang Bian, Jiefeng Long, Xiao Liang, Chao Feng, Hongsheng Li
arXiv:2607. 26621v2 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recommendation models (FRMs).
By Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou, Kuo Cai, Sheng Yu, Qiang Luo, Jian Liang, Ruiming Tang, Fei Pan, Peng Jiang, Wenwu Ou
arXiv:2607.26621v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendati...
By Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou, Kuo Cai, Sheng Yu, Qiang Luo, Jian Liang, Ruiming Tang, Fei Pan, Peng Jiang, Wenwu Ou
arXiv:2602.09839v2 Announce Type: replace
Abstract: Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional kno...
By Yijie Lin, Guofeng Ding, Haochen Zhou, Haobin Li, Mouxing Yang, Xi Peng
arXiv:2608. 10447v1 Announce Type: cross Abstract: Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predictions, often achieving higher accuracy than fast-thinking models that predict directly.
By Linh Dieu Le, Tong Chen, Shazia Sadiq, Hongzhi Yin, Ming Jin, Junliang Yu
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
By Junxian Wu, Chenghan Fu, Zhanheng Nie, Daoze Zhang, Bowen Wan, Wanxian Guan, Chuan Yu, Jian Xu, Bo Zheng