Hugging Face Trending Papers

Less is More: Encoder-only Audio-Visual Segmentation

Read the original on Hugging Face Trending Papers →

The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

Less is More: Encoder-only Audio-Visual Segmentation

The paper "Less is More: Encoder-only Audio-Visual Segmentation" introduces EASE, an encoder-only model for Audio‑Visual Semantic Segmentation (AVSS). EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—about three times faster than previous Transformer‑based AVSS models—and trains in under 11 GPU‑hours. The authors demonstrate that simpler, faster architectures can match or exceed the performance of more complex models across various backbones and resolutions.

By Ilpo Viertola, Vladimir Iashin, Sophie T\"otterstr\"om, Esa Rahtu
arXiv Computer Vision
Sep 3

Video Object Segmentation-Aware Audio Generation

The paper introduces a new task called video object segmentation‑aware audio generation, which conditions sound synthesis on object‑level segmentation maps. It presents SAGANet, a multimodal generative model that uses visual segmentation masks, video, and textual cues to produce controllable audio for musical instruments, offering fine‑grained, visually localized control. The authors also release the Segmented Music Solos dataset of instrument performance videos with segmentation information to support this task and demonstrate that SAGANet outperforms current state‑of‑the‑art methods in controllable, high‑fidelity Foley synthesis.

By Ilpo Viertola, Vladimir Iashin, Esa Rahtu
arXiv AI
Jul 2

LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter

arXiv:2607. 00687v1 Announce Type: cross Abstract: Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself.

By Tobias Christian Nauen, Anosh Billimoria, Federico Raue, Stanislav Frolov, Brian B. Moser, Andreas Dengel