arXiv Computer Vision

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

The paper introduces an interdisciplinary framework that uses AI and machine learning to analyze police body‑worn camera footage from the Rochester Police Department. It combines image, audio, and natural language processing—including speaker separation, transcription, and large language models—to detect and classify interaction patterns such as respect, disrespect, escalation, and de‑escalation. A custom evaluation pipeline assesses transcription quality and behavior detection accuracy, aiming to support law‑enforcement review, training, and accountability.

arXiv Computation and Language
Sep 11

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

The paper introduces BodyCam-VQA, an adaptive visual question answering framework designed to improve captioning of police body‑worn camera footage. By employing structured multimodal reasoning and probe question generation, the system extracts fine‑grained visual evidence that conventional captioning models miss. Experiments with various question generation models show that this VQA‑driven approach yields more reliable, objective, and detailed records of enforcement events.

By Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu
arXiv AI
Sep 4

Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling

The paper introduces a Short-Window Sliding Learning framework for real‑time violence detection in CCTV footage. It splits videos into 1–2 second clips and uses LLM‑based auto‑caption labeling to build fine‑grained datasets, preserving temporal continuity for accurate recognition of rapid violent events. The method achieves 95.25% accuracy on RWF‑2000 and 83.25% on UCF‑Crime, demonstrating strong generalization and real‑time applicability in intelligent surveillance systems.

By Seoik Jung, Taekyung Song, Yangro Lee, Sungjun Lee
arXiv Computer Vision
Aug 27

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.

By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
arXiv AI
Aug 11

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

arXiv:2608. 09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment.

By Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren
Hugging Face Trending Papers
Aug 11

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally.

Hugging Face Trending Papers
Jul 6

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.