arXiv AI By Yihua Zhang, Mingfu Liang, Jiyan Yang, Rong Jin, Wen-Yen Chen, Yiping Han, Huayu Li, Buyun Zhang, Liang Luo, Frank Shyu, Luke Simon, Sijia Liu, Tianlong Chen, Xi Liu

ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

Read the original on arXiv AI →

arXiv:2606. 28357v1 Announce Type: cross Abstract: Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reasoning and self-awareness of uncertainty.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 11

V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

V‑Retrver is an evidence‑driven retrieval framework that treats universal multimodal retrieval as an agentic reasoning process grounded in visual inspection. It allows multimodal large language models to selectively acquire visual evidence through external tools, alternating between hypothesis generation and targeted visual verification. The approach is trained with a curriculum that blends supervised activation, rejection‑based refinement, and reinforcement learning, achieving an average 23.0% improvement in retrieval accuracy across multiple benchmarks.

By Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan
arXiv Machine Learning
Sep 7

Latent-Aligned Reasoning for Multimodal Recommendation

The paper introduces LARK, a two‑stage latent reasoning framework designed to mitigate cross‑modal dilution in multimodal recommendation systems. In the first stage, learnable latent tokens are interleaved with chain‑of‑thought reasoning and aligned with a frozen vision encoder to preserve visual details. The second stage projects these latent representations through a bridge MLP, employing item‑to‑item contrastive learning and aligning intermediate features with the first‑stage hidden states to anchor final embeddings to the model’s reasoning output. Experiments on three public benchmarks and an industrial dataset demonstrate that LARK achieves state‑of‑the‑art performance across multiple recommendation architectures, with ablation studies confirming the contribution of each component.

By Jiarui Jin, Anyang Ji
arXiv AI
2d ago

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

arXiv:2610.01892v1 Announce Type: cross Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...

By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
arXiv AI
4d ago

ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.

By Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao