The paper introduces BodyCam-VQA, an adaptive visual question answering framework designed to improve captioning of police body‑worn camera footage. By employing structured multimodal reasoning and probe question generation, the system extracts fine‑grained visual evidence that conventional captioning models miss. Experiments with various question generation models show that this VQA‑driven approach yields more reliable, objective, and detailed records of enforcement events.
By Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu
arXiv:2606. 03686v1 Announce Type: new Abstract: We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent.
By Sarah Barrington, Maty Bohacek, Hany Farid
arXiv:2607.27245v2 Announce Type: replace-cross
Abstract: Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains...
By Vivek Senthil, Ernest Fokou\'e
The paper introduces a Short-Window Sliding Learning framework for real‑time violence detection in CCTV footage. It splits videos into 1–2 second clips and uses LLM‑based auto‑caption labeling to build fine‑grained datasets, preserving temporal continuity for accurate recognition of rapid violent events. The method achieves 95.25% accuracy on RWF‑2000 and 83.25% on UCF‑Crime, demonstrating strong generalization and real‑time applicability in intelligent surveillance systems.
By Seoik Jung, Taekyung Song, Yangro Lee, Sungjun Lee
arXiv:2608. 06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety.
By Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.
By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais