arXiv AI

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

arXiv AI
Aug 5

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

arXiv:2608. 02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains.

By Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
arXiv AI
Jun 17

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning

arXiv:2606. 17888v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxiliary signals, failing to capture the intricate and sample-specific dependencies between text and images in mathematical problem-solving.

By Wanshi Xu, Haokun Zhao, Haidong Yuan, Songjun Cao, Long Ma
arXiv AI
1d ago

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

SLVR (Structured Latent Visual Reasoning) is a training framework that blends explicit chain-of-thought with latent reasoning for multimodal large language models. It organizes reasoning into typed latent stages—planning, grounding, evidence selection, and integration—each supervised by corresponding signals such as plans, bounding boxes, visual evidence, and final rationales. Built on Qwen2.5‑VL‑7B, SLVR consistently improves performance on multimodal reasoning benchmarks, achieving significant gains on MMVP, BLINK Relation, V*, MathVista, and ChartQA.

By Albert Gao, Bing Xue, Andrea Zanette