arXiv Computer Vision By Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Jonathan Bateman, Adrian Martin, John McCluskey, Ernest Fokou\'e

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

Read the original on arXiv Computer Vision →

The paper introduces an interdisciplinary framework that uses AI and machine learning to analyze police body‑worn camera footage from the Rochester Police Department. It combines image, audio, and natural language processing—including speaker separation, transcription, and large language models—to detect and classify interaction patterns such as respect, disrespect, escalation, and de‑escalation. A custom evaluation pipeline assesses transcription quality and behavior detection accuracy, aiming to support law‑enforcement review, training, and accountability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 11

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

The paper introduces BodyCam-VQA, an adaptive visual question answering framework designed to improve captioning of police body‑worn camera footage. By employing structured multimodal reasoning and probe question generation, the system extracts fine‑grained visual evidence that conventional captioning models miss. Experiments with various question generation models show that this VQA‑driven approach yields more reliable, objective, and detailed records of enforcement events.

By Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu
arXiv AI
Sep 4

Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling

The paper introduces a Short-Window Sliding Learning framework for real‑time violence detection in CCTV footage. It splits videos into 1–2 second clips and uses LLM‑based auto‑caption labeling to build fine‑grained datasets, preserving temporal continuity for accurate recognition of rapid violent events. The method achieves 95.25% accuracy on RWF‑2000 and 83.25% on UCF‑Crime, demonstrating strong generalization and real‑time applicability in intelligent surveillance systems.

By Seoik Jung, Taekyung Song, Yangro Lee, Sungjun Lee
arXiv Computer Vision
Aug 27

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.

By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais