arXiv AI

(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

arXiv Computer Vision
Aug 27

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.

By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
arXiv AI
Aug 19

Convergent Evolution: How Different Language Models Learn Similar Number Representations

The paper shows that language models trained on natural text develop number representations that exhibit periodic features with dominant periods at T = 2, 5, 10. It identifies a two‑tiered hierarchy: all models learn Fourier‑domain spikes at these periods, but only some acquire geometrically separable features that allow linear classification of numbers modulo T. The study demonstrates that data, architecture, optimizer, and tokenizer influence whether these separable features emerge, and that models can learn them either from co‑occurrence signals in language or from multi‑token addition tasks, illustrating convergent evolution across diverse models.

By Deqing Fu, Tianyi Zhou, Mikhail Belkin, Vatsal Sharan, Robin Jia