arXiv AI By Ilpo Viertola, Vladimir Iashin, Sophie T\"otterstr\"om, Esa Rahtu

Less is More: Encoder-only Audio-Visual Segmentation

Read the original on arXiv AI →

The paper "Less is More: Encoder-only Audio-Visual Segmentation" introduces EASE, an encoder-only model for Audio‑Visual Semantic Segmentation (AVSS). EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—about three times faster than previous Transformer‑based AVSS models—and trains in under 11 GPU‑hours. The authors demonstrate that simpler, faster architectures can match or exceed the performance of more complex models across various backbones and resolutions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 24

Less is More: Encoder-only Audio-Visual Segmentation

The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.

arXiv AI
Jul 2

LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter

arXiv:2607. 00687v1 Announce Type: cross Abstract: Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself.

By Tobias Christian Nauen, Anosh Billimoria, Federico Raue, Stanislav Frolov, Brian B. Moser, Andreas Dengel