arXiv:2608.11645v2 Announce Type: replace
Abstract: Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame...
By Hossein Khalili (UCLA), Philip Do (UCLA), Alexander Vilesov (UCLA), Achuta Kadambi (UCLA), Kittipat Apicharttrisorn (Nokia Bell Labs), Nader Sehatbakhsh (UCLA)
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
By Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay
Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.
arXiv:2608.28691v1 Announce Type: cross
Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surround...
By Zhimin Li, Pan Wang, Jingxian Chen, Yuantao Tang, Anthony Chen, Qian Lou, Jingtong Hu
arXiv:2503.12232v3 Announce Type: replace
Abstract: Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensi...
By Yan Jiang, Hao Yu, Xu Cheng, Haoyu Chen, Zhaodong Sun, Guoying Zhao
arXiv:2608. 14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets.
By Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin