arXiv AI By Albert Gao, Bing Xue, Andrea Zanette

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

Read the original on arXiv AI →

SLVR (Structured Latent Visual Reasoning) is a training framework that blends explicit chain-of-thought with latent reasoning for multimodal large language models. It organizes reasoning into typed latent stages—planning, grounding, evidence selection, and integration—each supervised by corresponding signals such as plans, bounding boxes, visual evidence, and final rationales. Built on Qwen2.5‑VL‑7B, SLVR consistently improves performance on multimodal reasoning benchmarks, achieving significant gains on MMVP, BLINK Relation, V*, MathVista, and ChartQA.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

CoVA‑SFT is a new large‑scale dataset comprising 51.9K samples and over 222K multimodal reasoning steps that teach models to interleave text and visual abstractions across five layout families and 17 complex tasks. It includes explicit rationale formulations, agentic renderings, and verification loops to help models build and maintain internal visual workspaces for purely textual reasoning problems. A companion benchmark, CoVA‑Bench, contains 1,700 held‑out test samples for reproducible evaluation, and models fine‑tuned on CoVA‑SFT outperform all interleaved CoT baselines by more than 2× on average, though they still lag behind strong text‑only CoT baselines.

By Tsung-Han Wu, Heekyung Lee, Anya Ji, Haoming Chen, Trevor Darrell, Joseph E. Gonzalez, David M. Chan
arXiv Machine Learning
Aug 21

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

arXiv:2608. 19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage.

By Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi