Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper investigates whether detailed articulated human pose provides more discriminative power than coarse spatial relationships for early violence detection. By fixing the downstream pipeline and comparing five interaction representations—including bounding‑box geometry, handcrafted pose analogues, enriched pose descriptors, and a learned joint encoder—the study finds that pose‑based representations do not outperform coarse geometry. When visual encoders are frozen and evaluated on larger datasets, person‑crop appearance and whole‑frame context outperform geometry, but cropping to interacting people offers no advantage over encoding the entire frame. The authors further demonstrate that pre‑onset frames contain source‑related artifacts (e.g., title cards, watermarks) that contribute significantly to discrimination, suggesting that benchmark performance may reflect these artifacts rather than true event evidence.
arXiv:2606. 02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible.
Multi2AV‑Safety is a new benchmark for evaluating safety in multimodal-to-audio‑video generation. It covers all 11 non‑singleton conditioning configurations (text, image, audio, video) and contains 11,024 attack instances. The benchmark reveals that safety guards often fail when harmful semantics arise from combinations of benign inputs or when explicit harmful cues are masked by benign multimodal context, highlighting a gap in compositional risk perception.
arXiv:2604. 03329v2 Announce Type: replace-cross Abstract: Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible.
arXiv:2609.06991v1 Announce Type: cross Abstract: Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a single text prompt. This capability poses...
arXiv:2609.00206v1 Announce Type: cross Abstract: Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of...