arXiv Machine Learning

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

arXiv:2606. 05497v1 Announce Type: new Abstract: Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience.

arXiv Machine Learning
Jun 5

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.

By Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
arXiv AI
Sep 18

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

The paper introduces NeuroCognition, a benchmark based on three neuropsychological tests—Raven's Progressive Matrices, Spatial Working Memory, and the Wisconsin Card Sorting Test—to evaluate foundational cognitive abilities in large language models (LLMs). It finds that while LLMs excel on text tasks, their performance drops on image-based and more complex tasks, and they fail different parts of the same tasks compared to humans. NeuroCognition correlates with standard general-capability benchmarks yet measures distinct cognitive skills, highlighting where LLMs align with or diverge from human-like intelligence.

By Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh
arXiv AI
Aug 14

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.

By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv AI
Sep 7

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study

The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.

By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng
arXiv Computation and Language
4d ago

Cognitive Expert Language Models Better Align with the Corresponding Brain Systems

The study investigates whether language models tailored to specific cognitive domains better align with corresponding brain systems. By prompting and fine‑tuning large language models into six domain experts—sensory, spatial, numerical, reasoning, social, and abstract—the authors find that each expert’s representations more closely match the brain region associated with its domain than other experts. This domain‑specific alignment holds across multiple base models and fMRI datasets, while overall prediction accuracy remains largely unchanged, indicating that regional alignment can be obscured when summarizing across the brain.

By Zhivar Sourati, Mengxuan Helen Wu, Nona Ghazizadeh, Jonas Kaplan, Morteza Dehghani, Samuel A. Nastase
arXiv AI
Aug 18

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

arXiv:2608. 15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors.

By Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
arXiv Machine Learning
Sep 25

The Alignment Illusion in Multimodal Large Language Models

The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.

By Hong-Han Wang, Yuntao Wang, Hu Ding
arXiv Computation and Language
Sep 23

SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children

SpecialEduBench is a new benchmark for vision‑language models that evaluates their pedagogical competence in language intervention for autistic children across knowledge, skill, and attitude dimensions. It contains 4,537 knowledge items, 200 skill items, and 68 attitude items derived from recorded interventions, with 192 response cells for attitude items that combine pressure and monitoring. Eight leading vision‑language models were tested, none achieving full performance, especially in honesty cells, indicating that current models still struggle with situated teaching tasks.

By Jihoi Na, Taeyeong Kim, Sungjune Kong, Jaemin Jung, Min Joung Park, Kyungtae Joo, Ahhyun Kim, Shim Jaechang, Sooyoung Joo, Dongjin Ka, SeJoong Kim, Jimin Kim, HyunJin Jung, Unggi Lee
arXiv Computer Vision
6d ago

Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

arXiv:2609.31456v1 Announce Type: new Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...

By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy