Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.
arXiv:2605. 30794v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks.
arXiv:2508. 16129v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm.
arXiv:2606. 10833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored.
The paper introduces Cognitive Chain-of-Thought (CoCoT), a structured reasoning framework for vision‑language models that divides multimodal social reasoning into three cognitively inspired stages: Perception, Situation, and Norm. CoCoT improves performance across diverse tasks—multimodal intent disambiguation, theory of mind, social commonsense reasoning, and safety instruction following—by 5.9% to 4.6% on average. Fine‑tuning on CoCoT‑structured traces further boosts accuracy by 5–6% without explicit prompting, indicating that models internalize the structured reasoning pattern and that the approach enhances interpretability and social alignment in multimodal systems.
arXiv:2608.28623v1 Announce Type: cross Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before a...