What We are Missing in Multimodal LLM Evaluation?
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
arXiv:2507. 19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework.
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.
arXiv:2604. 18347v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) achieved rapid progress in the recent years.
arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.
arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.
arXiv:2506. 03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains.
arXiv:2602. 12279v2 Announce Type: replace-cross Abstract: Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their outputs.
arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.
arXiv:2503. 06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging.
arXiv:2608. 04586v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
arXiv:2603. 18558v2 Announce Type: replace-cross Abstract: Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows.
arXiv:2608. 04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
arXiv:2606. 07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues.