arXiv:2604. 20140v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex reasoning tasks.
By Darsh Kachroo, Arjun Prasaath Anbazhagan, Adriana Caraeni, Brennan Lagasse, Kevin Zhu
arXiv:2606. 03269v1 Announce Type: new Abstract: Visual Question Answering (VQA) is the task of answering questions about images, requiring the integration of multimodal input and reasoning.
By Thomas Eiter, Nelson Higuera Ruiz, Johannes Oetsch
arXiv:2607. 14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility.
By Abdullah Shaikh, Zain Naqi, Taha Zahid, Sandesh Kumar, Abdul Samad
The paper investigates whether reasoning always benefits universal multimodal embeddings (UMEs). By comparing the discriminative and reasoning-driven branches of UME-R1, the authors find that while reasoning improves positive similarity in 56.6% of cases, it also creates 15.7% false-helpful instances where hard negatives are drawn closer. Diagnostic analyses reveal that reasoning often de‑condenses retrieved neighborhoods and that chain‑of‑thought tokens encode evidence common to both positives and hard negatives. Based on these insights, the authors introduce SURE, a utility router that boosts UME-R1‑7B by 1.5 points and consistently improves other embedding models on MMEB‑V2 without retraining or extra VLM passes.
By Wenxiao Fan, Jingling Fu, Luohang Liu, Xinyuan Shan, Lichen Ma, Yu He, Junshi Huang, Yan Li, Kan Li
arXiv:2605. 28742v2 Announce Type: replace Abstract: Language models can use verifiable rewards to improve at a wide variety of reasoning tasks.
By Linas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan
MMEmb-R1 is a multimodal embedding framework that enhances reasoning by treating it as a latent variable and selecting beneficial reasoning paths through pair-aware selection and counterfactual intervention. It uses reinforcement learning to invoke reasoning only when necessary, reducing unnecessary computation and latency. On the MMEB-V2 benchmark, MMEmb-R1 achieves a state‑of‑the‑art score of 71.2 with just 4 B parameters.
By Yuchi Wang, Dingkang Yang, Haiyang Yu, Weikang Bian, Jiefeng Long, Xiao Liang, Chao Feng, Hongsheng Li