← Back to all news
arXiv AI September 2, 2026 By Zach Studdiford, Kanishka Misra

(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • rag
  • multimodal

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 9

Unveiling the Visual Counting Bottleneck in Vision-Language Models

arXiv:2605. 30170v2 Announce Type: replace-cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting.

By Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
llmsmultimodalbenchmarks
More like this →
arXiv AI
Jun 19

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

arXiv:2606. 19965v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts.

By Yihao Wang, Zijian He, Jie Ren, Keze Wang
llmsmultimodalbenchmarks
More like this →
arXiv Computation and Language
2d ago

Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

arXiv:2609.00293v1 Announce Type: new Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in contex...

By Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick
llmsragmultimodalsafety
More like this →
arXiv Machine Learning
Jun 8

Diagnosing Visual Ignorance in Vision-Language Models

arXiv:2606. 06890v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence.

By Runyu Zhou, Qi Zhang, Qixun Wang, Yisen Wang
llmsmultimodalbenchmarks
More like this →
arXiv AI
Jul 13

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.

By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
llmsfine-tuningmultimodalsafety
More like this →
arXiv AI
Jun 30

Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances

arXiv:2606. 29416v1 Announce Type: cross Abstract: Can a vision model truly see an object, or does it only fit surface-level visual cues?

By Xingyu Peng, Junran Wu, Yue Hou, Zhongliang Qiao, Jiaheng Liu, Shangzhe Li, Jichang Zhao, Wenjun Wu, Xianglong Liu, Yongxin Tong, Li Dong, Ke Xu
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea