arXiv Computer Vision By Haihan Li, Haihao Li, Zhenfei Xu, Jize Qian

Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

Read the original on arXiv Computer Vision →

The paper introduces CMPM, a Chinese Multi-Panel Meme benchmark comprising 1,214 annotated samples that capture five structural types, ordering dependencies, panel-order constraints, and optional comment context. It defines a two-layer evaluation: Task 1 tests structure typing and order-sensitive panel sequencing, while Task 2 assesses Chinese meme explanation generation using human ratings across visual, panel, humor, context, and faithfulness dimensions. Benchmarking five large vision‑language models shows that accuracy on canonical displays does not guarantee order understanding, as performance drops sharply under shuffled conditions, and that Gemini 3.1 Pro and GPT‑5.5 outperform open models in Task 2, with comment context providing only modest gains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 6

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

arXiv:2606. 05702v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored.

By Haoyu Zhou, Qing Qing, Caichong Li, Qixin Zhang, Yongcheng Jing, Ziqi Xu, Juncheng Hu, Xikun Zhang, Renqiang Luo
Hugging Face Trending Papers
Jun 4

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, we introduce a novel benchmark specifically designed to evaluate how VLMs perceive and reason about chronological information within and across images.

arXiv Computation and Language
3d ago

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.

By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick