MoVISA: Multi-Token Reasoning for Video Object Segmentation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.
arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Curre...
arXiv:2609.37042v1 Announce Type: cross Abstract: Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the lar...
arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
arXiv:2608.23234v1 Announce Type: new Abstract: In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language M...