arXiv AI By Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang

Depth-Guided Video Object Counting in Crowded Scenes

Read the original on arXiv AI →

arXiv:2608. 06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 17

DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting

DualCount introduces an instance-aware dual-decoder framework that couples density and point representations for zero‑shot object counting. By treating density estimation as a structured mass allocation over latent object instances, it applies two geometric constraints—per‑instance mass conservation and center‑of‑mass alignment—to enforce instance‑level consistency. Experiments on FSC‑147, PUCPR+, and CARPK demonstrate that this approach consistently reduces counting error and achieves new state‑of‑the‑art performance.

By Xuan Cuong Ngo
arXiv Computer Vision
Sep 17

Category Level 6D Object Pose Estimation from a Single RGB Image using Diffusion

The paper presents a generative framework that estimates category-level 6D pose and 3D size of objects from a single RGB image, using score-based diffusion models to produce a multi-hypothesis pose distribution. It replaces costly likelihood pruning with a Mean Shift approach to isolate the mode as the final pose estimate, achieving state-of-the-art results on the REAL275 benchmark. The method also decouples detection from pose estimation, enabling robust zero-shot generalisation on the Wild6D dataset and extending naturally to video sequences by propagating the pose distribution over time.

By Adam Bethell, Ravi Garg, Ian Reid
Hugging Face Trending Papers
Jul 8

HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting

Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such an environment are either private or not detailed per second.

arXiv AI
Jul 9

HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting

arXiv:2607. 07322v1 Announce Type: cross Abstract: Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people.

By Reem AlYabis, Fares AlTuwaim, AlJawharh AlOtaibi, Mohamed Eltahir