Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Related stories
When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?
arXiv:2608.22174v1 Announce Type: new Abstract: Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding...
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models
arXiv:2608. 11907v1 Announce Type: cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
arXiv:2508. 04427v2 Announce Type: replace-cross Abstract: Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks.
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs).
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.
A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hin...
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.
SenseNova-U1.5: Towards Native Unified Visual Intelligence
SenseNova-U1.5 is an 8B‑MoT native unified multimodal model that can understand, reason about, and generate visual content without using an encoder or VAE. It improves visual fidelity and text rendering through spatially coherent patch reconstruction, large‑scale training on curated generation and editing data, and native resolutions up to 4K. Post‑training, specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing are optimized and distilled into a multi‑expert framework, yielding advances in image fidelity, complex composition, multi‑reference editing, and instruction following.
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
arXiv:2508.13680v5 Announce Type: replace-cross Abstract: We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (V...