arXiv AI

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

arXiv:2608. 13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels.

arXiv AI
Jun 26

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

arXiv:2606. 26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence.

By Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv