arXiv Computer Vision By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

Read the original on arXiv Computer Vision →

DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv Computer Vision
1d ago

FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning

FOMO is a training‑based selective video unlearning method that prioritizes preserving the original scene while removing targeted concepts. It localizes concept‑related representations for modification and employs a preservation mechanism that maintains non‑target scene information without auxiliary data. The approach extends to motion unlearning, enabling removal of concepts defined by temporal behavior, and achieves a strong balance between concept removal and scene preservation.

By {\L}ukasz Rudnik, Agnieszka Polowczyk, Alicja Polowczyk, Przemys{\l}aw Spurek