arXiv Computer Vision By Dayoung Kil, Seong-heum Kim

Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

Read the original on arXiv Computer Vision →

Efficient Unified Multimodal Understanding (EUMU) is the winning solution for the MUMU Track of the 8th LSVOS Challenge, addressing multi‑concept image tagging, open‑vocabulary object detection, and image captioning with a single efficient model. It leverages a shared pretrained multimodal backbone and lightweight heads, while applying task‑aware inference refinement that uses detection cues to improve captioning, caption cues to recover missed detections, and image statistics to refine tagging. With 239.169 M parameters, 23.947 GFLOPs, and 4.5 GB peak memory, EUMU achieves a challenge score of 17.3409 and is publicly available on GitHub.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 17

CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026

The paper presents CapMap-MS-TTA, a unified multimodal model that achieved third place in the MUMU track of the 8th Large-scale Video Object Segmentation Challenge. The solution tackles image tagging, open‑vocabulary object detection, and English captioning under strict resource limits, using an expanded keyword lexicon and lightweight expand‑hints for tagging, and Florence‑2 with multi‑scale and horizontal‑flip TTA plus label‑aware NMS for detection. The approach improves the reproduced Florence‑2 baseline from 15.16 to a best public score of 16.4815 without fine‑tuning.

By Chengfeng Qiu, Kaifeng Wei
arXiv Computer Vision
Sep 2

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

arXiv:2609.00591v1 Announce Type: new Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...

By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon
Hugging Face Trending Papers
Aug 12

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs).