The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.
By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and...
ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.
By Anindya Mondal, Sauradip Nag, Anjan Dutta
arXiv:2607. 06420v1 Announce Type: cross Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning.
By Jinhong Deng, Limeng Qiao, Guanglu Wan
ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is built on existing 3B-parameter unified foundation model and is adapted for object localization tasks using three key innovations: density-aware adaptive zooming with objectness maps for spatial grounding; a boundary-aware count policy via GRPO to eliminate crop-boundary errors; and a cycle-consistent GRPO strategy where the understanding branch self-critiques generated outputs, closing the understanding-generation gap without any external annotations.
arXiv:2609.17613v1 Announce Type: new
Abstract: Zero-shot object counting aims to estimate the number of objects specified by a text query without category-specific training. Recent approaches primar...
By Xuan Cuong Ngo
arXiv:2608.20621v1 Announce Type: new
Abstract: Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coa...
By Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh
arXiv:2508. 10956v3 Announce Type: replace-cross Abstract: Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying and recognizing low-level details and higher-level abstractions.
By Abhishek Kolari, Mohammadhossein Khojasteh, Yifan Jiang, Floris den Hengst, Filip Ilievski
NumBench is a large benchmark of 640,000 text‑to‑image prompts that tests how well models count objects, covering 1,600 categories and counts from 1 to 100. It uses a factorial design to vary composition, spatial guidance, and appearance while balancing counts, and introduces a process model that predicts a near‑quadratic collision deficit at low occupancy. The authors also propose the Confidence‑Weighted Numeric Precision Score for scalable evaluation and find that all tested systems perform poorly above 50 objects, with count range, layout, and composition having the largest effects.
By Sandeep Wadhwa, Mayank Vatsa, Richa Singh, Parrva Chirag Shah, Prakhar Galriya
arXiv:2607. 09544v1 Announce Type: cross Abstract: Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting.
By Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov, Kentaro Inui
arXiv:2609.13706v1 Announce Type: new
Abstract: Co-salient object detection (Co-SOD) requires a model to find foreground regions that are salient in individual images and supported by the image group...
By Yuan Xiang, Matteo Rossi, Yingzhou Chen
arXiv:2605. 30170v2 Announce Type: replace-cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting.
By Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan