arXiv Machine Learning By Dongchen Lu, Zhimo Li, Mao Shu, Huo Cao

DeepLatent: Think with Images via Parallel Latent Visual Reasoning

Read the original on arXiv Machine Learning →

arXiv:2606. 00562v1 Announce Type: cross Abstract: The emerging paradigm of "thinking with images" embeds visual states into intermediate reasoning steps, defining a new frontier for Vision-Language Models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 17

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning

arXiv:2606. 17888v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxiliary signals, failing to capture the intricate and sample-specific dependencies between text and images in mathematical problem-solving.

By Wanshi Xu, Haokun Zhao, Haidong Yuan, Songjun Cao, Long Ma