arXiv Computer Vision

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

HakushoBench is a Japanese chart and table visual question answering benchmark created from 33 governmental white papers, comprising 2,053 images across more than ten types. The dataset includes manually annotated and independently verified QA pairs that evaluate holistic understanding of charts and tables rather than just local visual cues. Experiments show that HakushoBench is significantly harder than existing Japanese benchmarks, with sub‑10B open‑weight models achieving at most 58.6% accuracy and even large models like Qwen3.5-397B-A17B lagging behind Gemini‑3‑Pro by 8.1 points.

arXiv AI
Aug 19

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

BEAR-Bench is a bilingual benchmark for multimodal large language models, featuring 1,000 human‑annotated questions derived from text‑rich business and scientific documents in English and Russian. It evaluates 16 MLLMs, including Gemini 3.1 Pro and Qwen3.5‑397B, revealing significant performance gaps even for the strongest systems. The benchmark also serves to compare hallucination‑detection methods by analyzing model failures on these complex documents.

By Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
arXiv AI
Sep 2

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

SCAFFOLD is a large-scale structured dataset of computer science research figures, each paired with captions, context, questions, answers, and chain-of-thought reasoning traces. It contains 157,387 figure–question pairs from 3,058 arXiv papers, with subsets of 36,797 and 12,000 pairs for medium and small-scale use. The dataset was created using layout detection, PDF parsing, and AI-assisted question generation, and was used to benchmark a vision‑language model (Qwen2.5‑VL‑3B‑Instruct).

By Ranjit Raut, Aarav Subedi, Sagun Rai, Sudan Jha
arXiv AI
Jul 8

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv:2607. 05614v1 Announce Type: cross Abstract: Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications.

By Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha, Ezharuddin Jubaer, Armun Alam, Pranjal Kumar Nandi, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman
arXiv Computer Vision
Aug 31

CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models

The paper introduces CompareBench, a new benchmark suite for evaluating visual comparison reasoning in vision‑language models. It includes TallyBench for object counting, OmniCaps for captioning and tagging, and a 1,200‑question CompareBench that tests quantity, geometric, spatial, and temporal comparisons. Experiments on nine closed‑source models show strong overall performance but persistent weaknesses in counting, spatial reasoning, geometric comparison, and temporal ordering, highlighting visual comparison as a systematic challenge for current VLMs.

By Jie Cai