arXiv AI By Damith Chamalke Senadeera, Dimitrios Kollias, Gregory Slabaugh

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

Read the original on arXiv AI →

arXiv:2604. 03329v2 Announce Type: replace-cross Abstract: Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 3

From Visual Cues to Spoken Narration: Rethinking Audio Description

The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.

By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach
arXiv AI
Jul 16

A Hybrid Mamba for Audio-Visual Navigation

arXiv:2607. 13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences.

By Yi Wang, Yinfeng Yu
arXiv Computer Vision
4d ago

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

arXiv:2604.15086v3 Announce Type: replace-cross Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...

By Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan