Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces BodyCam-VQA, an adaptive visual question answering framework designed to improve captioning of police body‑worn camera footage. By employing structured multimodal reasoning and probe question generation, the system extracts fine‑grained visual evidence that conventional captioning models miss. Experiments with various question generation models show that this VQA‑driven approach yields more reliable, objective, and detailed records of enforcement events.
The paper introduces an interdisciplinary framework that uses AI and machine learning to analyze police body‑worn camera footage from the Rochester Police Department. It combines image, audio, and natural language processing—including speaker separation, transcription, and large language models—to detect and classify interaction patterns such as respect, disrespect, escalation, and de‑escalation. A custom evaluation pipeline assesses transcription quality and behavior detection accuracy, aiming to support law‑enforcement review, training, and accountability.
The paper introduces an audio-first triage method for budgeted vision‑language captioning of untrimmed egocentric video. By selecting windows for a vision‑language model based on lightweight audio features before any video decoding, the approach reduces costly model calls. It achieves 4.0–10.8 percentage point improvements in action coverage across call rates, cuts 9–20% of VLM calls at matched coverage on EPIC‑KITCHENS‑100, and outperforms uniform sampling and recent visual keyframe selectors on Ego4D.
arXiv:2609.06991v1 Announce Type: cross Abstract: Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a single text prompt. This capability poses...
Video-HolmesV2 is a new benchmark that tests multimodal large language models on their ability to reason with spatio‑temporal audio‑visual evidence in long videos. It requires models to justify answers with precise evidence, uses a multi‑model cross‑verification pipeline and a spatio‑temporal evidence‑aware metric, and introduces an audio‑text guided token compression framework to reduce long‑context noise. In evaluations, even strong proprietary models score below 60% while the proposed approach outperforms comparable open‑source omni‑models.
arXiv:2606. 05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement.