arXiv:2607. 07322v1 Announce Type: cross Abstract: Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people.
By Reem AlYabis, Fares AlTuwaim, AlJawharh AlOtaibi, Mohamed Eltahir
arXiv:2609.23012v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene...
By Moshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush, Reda Bendraou
arXiv:2608. 06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts.
By Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
arXiv:2403. 09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation.
By Yiming Ma, Victor Sanchez, Tanaya Guha
ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.
By Anindya Mondal, Sauradip Nag, Anjan Dutta
DualCount introduces an instance-aware dual-decoder framework that couples density and point representations for zero‑shot object counting. By treating density estimation as a structured mass allocation over latent object instances, it applies two geometric constraints—per‑instance mass conservation and center‑of‑mass alignment—to enforce instance‑level consistency. Experiments on FSC‑147, PUCPR+, and CARPK demonstrate that this approach consistently reduces counting error and achieves new state‑of‑the‑art performance.
By Xuan Cuong Ngo