arXiv AI By Yuying Li, Siyi Qian, Hao Liang, Leqi Zheng, Ruichuan An, Linzhuang Sun, Jiajun Zhang, Wentao Zhang

CapGeo-Bench: Decoupling Visual Perception from Reasoning and Evaluating Geometric Understanding

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computer Vision
Aug 31

CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models

The paper introduces CompareBench, a new benchmark suite for evaluating visual comparison reasoning in vision‑language models. It includes TallyBench for object counting, OmniCaps for captioning and tagging, and a 1,200‑question CompareBench that tests quantity, geometric, spatial, and temporal comparisons. Experiments on nine closed‑source models show strong overall performance but persistent weaknesses in counting, spatial reasoning, geometric comparison, and temporal ordering, highlighting visual comparison as a systematic challenge for current VLMs.

By Jie Cai
arXiv Computer Vision
Sep 14

ChitraMiti: Benchmarking Visual Grounding and Modality Reliance in Bengali Geometric Reasoning

ChitraMiti introduces a synthetic benchmark of 12,874 Bengali planar geometry problems with structured 15‑attribute descriptions, alongside a complementary set of 500 textbook diagrams. Using a three‑phase protocol that tests diagram‑only, diagram‑plus‑description, and description‑only inputs, the study finds that description‑only performance matches diagram‑plus‑description performance across several VLMs, yet models still struggle with cross‑modal verification. Fine‑tuning on ChitraMiti improves results on both datasets, though a gap remains compared to the best zero‑shot model.

By Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Sumaiya Tabassum Nimi, Md Adnan Arefeen
arXiv Computation and Language
Sep 10

From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

The paper introduces a framework that combines a Geometric Vision Parser and a Symbolic Solver to enable a Large Language Model to solve complex plane geometry problems. By translating diagrams into symbolic representations and performing formal deductions, the approach reduces hallucinations and produces interpretable, human-like solutions. Experiments on a new benchmark from 2025 Chinese Zhongkao exams show performance comparable to Gemini 2.5 Pro.

By Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao, Xin Shen, Dongcai Lu, Yi Zhou
arXiv AI
Aug 5

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

arXiv:2607. 27670v2 Announce Type: replace-cross Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions.

By Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao
arXiv Computation and Language
3d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong