arXiv AI By Chengfeng Qiu, Kaifeng Wei

CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026

Read the original on arXiv AI →

The paper presents CapMap-MS-TTA, a unified multimodal model that achieved third place in the MUMU track of the 8th Large-scale Video Object Segmentation Challenge. The solution tackles image tagging, open‑vocabulary object detection, and English captioning under strict resource limits, using an expanded keyword lexicon and lightweight expand‑hints for tagging, and Florence‑2 with multi‑scale and horizontal‑flip TTA plus label‑aware NMS for detection. The approach improves the reproduced Florence‑2 baseline from 15.16 to a best public score of 16.4815 without fine‑tuning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 18

Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

Efficient Unified Multimodal Understanding (EUMU) is the winning solution for the MUMU Track of the 8th LSVOS Challenge, addressing multi‑concept image tagging, open‑vocabulary object detection, and image captioning with a single efficient model. It leverages a shared pretrained multimodal backbone and lightweight heads, while applying task‑aware inference refinement that uses detection cues to improve captioning, caption cues to recover missed detections, and image statistics to refine tagging. With 239.169 M parameters, 23.947 GFLOPs, and 4.5 GB peak memory, EUMU achieves a challenge score of 17.3409 and is publicly available on GitHub.

By Dayoung Kil, Seong-heum Kim
arXiv AI
Sep 21

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.

By Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim
Hugging Face Trending Papers
Jul 23

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.

arXiv Machine Learning
Jun 30

DataComp-VLM: Improved Open Datasets for Vision-Language Models

arXiv:2606. 28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies.

By Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy