← Back to all news
arXiv Computer Vision September 30, 2026 By Liwei Che, Yihao Quan, Sen Fang, Hongyi Wang, Ranjay Krishna, Ruixiang Tang, Vladimir Pavlovic

Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • fine-tuning
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 7

HoloCount: A Holistic Visual Counting Benchmark for MLLMs

arXiv:2607. 06420v1 Announce Type: cross Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning.

By Jinhong Deng, Limeng Qiao, Guanglu Wan
llmsmultimodalbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 9

Unveiling the Visual Counting Bottleneck in Vision-Language Models

arXiv:2605. 30170v2 Announce Type: replace-cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting.

By Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
llmsmultimodalbenchmarks
More like this →
arXiv Machine Learning
Jul 13

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

arXiv:2607. 09544v1 Announce Type: cross Abstract: Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting.

By Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov, Kentaro Inui
llmsmultimodal
More like this →
arXiv AI
Jun 3

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.

By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
llmscomputer-visionmultimodalbenchmarkssafety
More like this →
arXiv Computer Vision
1d ago

Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores

arXiv:2605.12491v2 Announce Type: replace Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...

By Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo
llmscomputer-visionbenchmarks
More like this →
arXiv Computer Vision
Aug 24

RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars

arXiv:2608.20621v1 Announce Type: new Abstract: Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coa...

By Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh
diffusionbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea