arXiv Machine Learning By Kazuma Kano, Yuki Mori, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi

CorVS+: Correspondence-Driven Association of Video Trajectories and Sensors for Identity-Aware Person Localization in Warehouses

Read the original on arXiv Machine Learning →

arXiv:2510. 26369v2 Announce Type: replace Abstract: Logistics warehouses have struggled with labor shortages, but the inbound processes remain particularly human-powered.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 8

Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking

arXiv:2606. 07233v1 Announce Type: cross Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments.

By Eduardo Borges, Lu\'is Garrote, Urbano J. Nunes
Hugging Face Trending Papers
Jul 9

Whareformer: Learning to Track What is Where in Long Egocentric Videos

The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become heavily occluded. In this paper, we propose the first learning-based solution to the OSNOM task: Whareformer, a transformer-based model with two components: an updatable memory of established tracks and a track assignment module that associates observations with existing tracks in a feed-forward manner.

arXiv Computer Vision
4d ago

Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces

Physical AI Smart Spaces is the first benchmark that offers large‑scale, multi‑class, multi‑camera 3D perception data for indoor smart spaces. It includes over 280 hours of synchronized 1080p footage from nearly 1,800 cameras in warehouses, hospitals, and retail venues, with automatic annotations for identities, 2D and 3D bounding boxes, camera calibration, and depth. The benchmark spans synthetic generation, appearance augmentation, and real‑world Sim2Real evaluation, and introduces a 3D version of Higher Order Tracking Accuracy (HOTA) for evaluating multi‑class 3D box tracking.

By Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang
arXiv Computer Vision
Sep 1

Everybody Tracking Every Body

arXiv:2608.29927v1 Announce Type: new Abstract: We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each indivi...

By Daeyun Shin, Yunhan Zhao, Shu Kong, Alexander C. Berg, Charless Fowlkes
arXiv Computer Vision
Sep 7

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

The paper introduces MovingDroneCrowd++, a large-scale video dataset for dense crowd counting and tracking from moving drones, featuring varied flight altitudes, camera angles, and lighting. It presents two new methods: GD3A for Video Individual Counting and GIA-Track for Multi-Object Tracking, both leveraging group-wise density assignment and identity association to handle aerial challenges. Experiments demonstrate significant improvements, reducing counting error by 47.4% and boosting tracking accuracy by 64.6%.

By Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
arXiv AI
Aug 12

TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification

arXiv:2504. 11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual surveys, Bluetooth/WiFi tracking, and Automated Passenger Counters, are often costly, device-dependent, or unable to support individual-level matching.

By Kaicong Huang, Talha Azfar, Jack Reilly, Ruimin Ke