arXiv AI By Rohit Sinha, Kunal Tilaganji, Tanuja Ganu, Nagarajan Natarajan, Amit Sharma, Vineeth Balasubramanian

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

Read the original on arXiv AI →

arXiv:2608. 10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

arXiv:2607. 10966v1 Announce Type: new Abstract: We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning.

By Mingyuan Wu, Jingcheng Yang, Shengyi Qian, Xudong Wang, Jize Jiang, Qifan Wang, Aashu Singh, Khoi Pham, Fei Liu, Zhaolun Su, Zhuokai Zhao, Klara Nahrstedt, Jianyu Wang, Hanchao Yu
arXiv AI
Sep 21

Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation

CRYSTAL is a diagnostic benchmark comprising 6,372 multimodal reasoning instances that assess models through verifiable intermediate steps. It introduces two metrics—Match F1 and Ordered Match F1—to evaluate step-level precision, recall, and order. The benchmark, built via a Delphi-inspired pipeline with four independent MLLMs, reveals systematic failures in current models, such as cherry‑picking and disordered reasoning, and proposes a Causal Process Reward and CPR‑Curriculum to improve reasoning performance.

By Wayner Barrios, SouYoung Jin