arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.
By Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
arXiv:2510.01030v2 Announce Type: replace
Abstract: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust repres...
By Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh
The paper introduces NeuroCognition, a benchmark based on three neuropsychological tests—Raven's Progressive Matrices, Spatial Working Memory, and the Wisconsin Card Sorting Test—to evaluate foundational cognitive abilities in large language models (LLMs). It finds that while LLMs excel on text tasks, their performance drops on image-based and more complex tasks, and they fail different parts of the same tasks compared to humans. NeuroCognition correlates with standard general-capability benchmarks yet measures distinct cognitive skills, highlighting where LLMs align with or diverge from human-like intelligence.
By Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh
arXiv:2607. 24999v1 Announce Type: cross Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them.
By Dengzhe Hou, Lingyu Jiang, Fangzhou Lin, Kazunori D Yamada
arXiv:2511. 16107v3 Announce Type: replace-cross Abstract: Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training.
By Shao-Jun Xia, Huixin Zhang, Zhengzhong Tu
arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.
By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.
By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng
The study investigates whether language models tailored to specific cognitive domains better align with corresponding brain systems. By prompting and fine‑tuning large language models into six domain experts—sensory, spatial, numerical, reasoning, social, and abstract—the authors find that each expert’s representations more closely match the brain region associated with its domain than other experts. This domain‑specific alignment holds across multiple base models and fMRI datasets, while overall prediction accuracy remains largely unchanged, indicating that regional alignment can be obscured when summarizing across the brain.
By Zhivar Sourati, Mengxuan Helen Wu, Nona Ghazizadeh, Jonas Kaplan, Morteza Dehghani, Samuel A. Nastase
arXiv:2608. 15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors.
By Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.
By Hong-Han Wang, Yuntao Wang, Hu Ding
SpecialEduBench is a new benchmark for vision‑language models that evaluates their pedagogical competence in language intervention for autistic children across knowledge, skill, and attitude dimensions. It contains 4,537 knowledge items, 200 skill items, and 68 attitude items derived from recorded interventions, with 192 response cells for attitude items that combine pressure and monitoring. Eight leading vision‑language models were tested, none achieving full performance, especially in honesty cells, indicating that current models still struggle with situated teaching tasks.
By Jihoi Na, Taeyeong Kim, Sungjune Kong, Jaemin Jung, Min Joung Park, Kyungtae Joo, Ahhyun Kim, Shim Jaechang, Sooyoung Joo, Dongjin Ka, SeJoong Kim, Jimin Kim, HyunJin Jung, Unggi Lee
arXiv:2609.31456v1 Announce Type: new
Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...
By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy