The paper presents CapMap-MS-TTA, a unified multimodal model that achieved third place in the MUMU track of the 8th Large-scale Video Object Segmentation Challenge. The solution tackles image tagging, open‑vocabulary object detection, and English captioning under strict resource limits, using an expanded keyword lexicon and lightweight expand‑hints for tagging, and Florence‑2 with multi‑scale and horizontal‑flip TTA plus label‑aware NMS for detection. The approach improves the reproduced Florence‑2 baseline from 15.16 to a best public score of 16.4815 without fine‑tuning.
By Chengfeng Qiu, Kaifeng Wei
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
arXiv:2608. 11907v1 Announce Type: cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
arXiv:2609.00591v1 Announce Type: new
Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...
By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs).
arXiv:2603.14733v2 Announce Type: replace
Abstract: Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos...
By Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian, Xinya Du, Yunhui Guo, Vibhav Gogate