ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.
By Anindya Mondal, Sauradip Nag, Anjan Dutta
The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.
By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and...
arXiv:2508.16644v5 Announce Type: replace
Abstract: Diffusion models excel at photorealistic synthesis but struggle with object count fidelity, especially in high-density settings. We introduce COUNT...
By Anindya Mondal, Sauradip Nag, Ayan Banerjee, Josep Llados, Xiatian Zhu, Anjan Dutta
DualCount introduces an instance-aware dual-decoder framework that couples density and point representations for zero‑shot object counting. By treating density estimation as a structured mass allocation over latent object instances, it applies two geometric constraints—per‑instance mass conservation and center‑of‑mass alignment—to enforce instance‑level consistency. Experiments on FSC‑147, PUCPR+, and CARPK demonstrate that this approach consistently reduces counting error and achieves new state‑of‑the‑art performance.
By Xuan Cuong Ngo
The paper introduces Group-Individual Object Counting (GIC), a new task that requires models to count both individual objects and higher‑level semantic groups within the same image. To support this, the authors present BunchCount, a benchmark of 1,330 images with 89,254 individual and 11,065 group annotations that include explicit containment relations. Experiments show that existing counting models excel at individual counting but struggle with group counting, leading the authors to propose a relational counting framework that leverages group‑individual containment to improve group‑level accuracy while preserving individual performance.
By Rui Wang, Junyi Huang, Jiahui Li, Qiao Yu, Yixue Hao, Long Hu, Baoru Huang
arXiv:2609.37096v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. I...
By Liwei Che, Yihao Quan, Sen Fang, Hongyi Wang, Ranjay Krishna, Ruixiang Tang, Vladimir Pavlovic
arXiv:2501.04001v4 Announce Type: replace
Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...
By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained VLMs through direct QA-style fine-tuning from a single global image. We argue that this direct paradigm remains limited for in-the-wild deployment in terms of operational range, reliability under reduced-resolution inputs, and inference efficiency.
arXiv:2608.20621v1 Announce Type: new
Abstract: Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coa...
By Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh
The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.
By Zongjian Wu, Lei Zhang
RankGround is a two‑stage framework for GUI grounding that uses a single Vision‑Language Model call per query. It introduces GroundRanker, a lightweight multimodal reranker that selects the most promising crop from a dense candidate set, trained with a two‑stage curriculum on ranking supervision data derived from existing grounding datasets. Experiments show RankGround outperforms strong baselines, achieving 1.4× faster inference and a 5.5% average improvement in localization accuracy over the second‑best method across all backbones and screen scales.
By Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang