Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such an environment are either private or not detailed per second.
arXiv:2609.23012v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene...
By Moshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush, Reda Bendraou
arXiv:2403. 09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation.
By Yiming Ma, Victor Sanchez, Tanaya Guha
arXiv:2608. 06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts.
By Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.
By Anindya Mondal, Sauradip Nag, Anjan Dutta
arXiv:2609.09396v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...
By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
DualCount introduces an instance-aware dual-decoder framework that couples density and point representations for zero‑shot object counting. By treating density estimation as a structured mass allocation over latent object instances, it applies two geometric constraints—per‑instance mass conservation and center‑of‑mass alignment—to enforce instance‑level consistency. Experiments on FSC‑147, PUCPR+, and CARPK demonstrate that this approach consistently reduces counting error and achieves new state‑of‑the‑art performance.
By Xuan Cuong Ngo
The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.
By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.
By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
The paper introduces Group-Individual Object Counting (GIC), a new task that requires models to count both individual objects and higher‑level semantic groups within the same image. To support this, the authors present BunchCount, a benchmark of 1,330 images with 89,254 individual and 11,065 group annotations that include explicit containment relations. Experiments show that existing counting models excel at individual counting but struggle with group counting, leading the authors to propose a relational counting framework that leverages group‑individual containment to improve group‑level accuracy while preserving individual performance.
By Rui Wang, Junyi Huang, Jiahui Li, Qiao Yu, Yixue Hao, Long Hu, Baoru Huang
NumBench is a large benchmark of 640,000 text‑to‑image prompts that tests how well models count objects, covering 1,600 categories and counts from 1 to 100. It uses a factorial design to vary composition, spatial guidance, and appearance while balancing counts, and introduces a process model that predicts a near‑quadratic collision deficit at low occupancy. The authors also propose the Confidence‑Weighted Numeric Precision Score for scalable evaluation and find that all tested systems perform poorly above 50 objects, with count range, layout, and composition having the largest effects.
By Sandeep Wadhwa, Mayank Vatsa, Richa Singh, Parrva Chirag Shah, Prakhar Galriya
arXiv:2608.20621v1 Announce Type: new
Abstract: Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coa...
By Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh