arXiv AI

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

arXiv:2604. 03329v2 Announce Type: replace-cross Abstract: Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible.

arXiv AI
Jul 16

A Hybrid Mamba for Audio-Visual Navigation

arXiv:2607. 13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences.

By Yi Wang, Yinfeng Yu
Hugging Face Trending Papers
Aug 6

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.