arXiv Machine Learning By Andr\'e G. Viveiros, Nuno Gon\c{c}alves, Matthias Lindemann, Andr\'e Martins

LanteRn: Latent Visual Structured Reasoning

Read the original on arXiv Machine Learning →

arXiv:2603. 25629v2 Announce Type: replace-cross Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
5d ago

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.

By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
arXiv Machine Learning
Sep 7

Latent-Aligned Reasoning for Multimodal Recommendation

The paper introduces LARK, a two‑stage latent reasoning framework designed to mitigate cross‑modal dilution in multimodal recommendation systems. In the first stage, learnable latent tokens are interleaved with chain‑of‑thought reasoning and aligned with a frozen vision encoder to preserve visual details. The second stage projects these latent representations through a bridge MLP, employing item‑to‑item contrastive learning and aligning intermediate features with the first‑stage hidden states to anchor final embeddings to the model’s reasoning output. Experiments on three public benchmarks and an industrial dataset demonstrate that LARK achieves state‑of‑the‑art performance across multiple recommendation architectures, with ablation studies confirming the contribution of each component.

By Jiarui Jin, Anyang Ji
arXiv Machine Learning
Sep 18

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

Uni-LaDiR (Unified Latent Diffusion Reasoner) is a new framework that unifies multimodal reasoning by mapping teacher reasoning steps from different modalities into a shared latent space of thought tokens. It employs a diffusion model to predict the next block of thought tokens, jointly training the encoder and reasoner with shared weights to ensure tokens are both useful and predictable. The approach achieves relative gains of 7.3% on visual reasoning benchmarks and 6.1% on robot manipulation tasks compared to the strongest baselines.

By Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Yian Ma, Lianhui Qin