arXiv:2607. 07322v1 Announce Type: cross Abstract: Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people.
By Reem AlYabis, Fares AlTuwaim, AlJawharh AlOtaibi, Mohamed Eltahir
Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such an environment are either private or not detailed per second.
The paper explores the use of Closed‑Circuit Television (CCTV) footage to estimate urban rail platform crowding in real time. It compares three computer‑vision methods—object detection and counting, crowd‑level classification with a Vision Transformer, and semantic segmentation—to extract crowd-related features. A novel convex ridge regression technique is introduced to convert segmentation outputs into passenger counts, and the methods are evaluated on a privacy‑preserving dataset of over 600 hours of Washington Metropolitan Area Transit Authority (WMATA) video, showing that CCTV alone can provide valuable real‑time crowd estimates.
By Riccardo Fiorista, Awad Abdelhalim, Anson F. Stewart, Gabriel L. Pincus, Ian Thistle, Jinhua Zhao
arXiv:2608. 06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts.
By Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.
By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.
By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.
By Anindya Mondal, Sauradip Nag, Anjan Dutta
Crowd counting is a fundamental task in computer vision. However, crowd counting in low-light environments remains largely underexplored, despite its practical importance in the real world.
arXiv:2609.01027v1 Announce Type: new
Abstract: Out-of-distribution (OOD) detection predicts whether a test image belongs to none of the predefined classes. To evaluate this task, benchmarks need ima...
By Ruslan Rozumnyi, Mat\v{e}j Such\'anek, Tom\'a\v{s} Voj\'i\v{r}, Kl\'ara Janou\v{s}kov\'a, Ji\v{r}\'i Matas
arXiv:2609.23012v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene...
By Moshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush, Reda Bendraou
arXiv:2608.31107v1 Announce Type: new
Abstract: The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalizatio...
By Lucas Wojcik, Gabriel E. Lima, Sergio M. Silva Jr., Eduil Nascimento Jr., David Menotti
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and...