arXiv:2609.36569v1 Announce Type: cross
Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...
By Yupeng Chang, Wenxuan Zhang, Yuan Wu
arXiv:2609.39934v1 Announce Type: cross
Abstract: Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable proba...
By Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan
arXiv:2609.37169v1 Announce Type: cross
Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since add...
By Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou
arXiv:2606. 26079v1 Announce Type: cross Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines.
By Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
arXiv:2602. 09689v2 Announce Type: replace Abstract: Fine-tuning large pre-trained models on a target distribution often improves in-distribution (ID) accuracy, but at the cost of out-of-distribution (OOD) robustness as representations specialize to the fine-tuning data.
By Alireza Abdollahpoorrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal Frossard
The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.
By Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu
arXiv:2607. 15094v1 Announce Type: cross Abstract: Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases.
By Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu
arXiv:2603. 00546v2 Announce Type: replace Abstract: Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains.
By Zeyu Chen, Huanjin Yao, Ziwang Zhao, Min Yang
The paper demonstrates that multiple‑choice visual question answering (MC‑VQA) benchmarks are unreliable because model performance is highly sensitive to semantically neutral prompt formatting choices—such as option ID sets, delimiters, and separators—despite protocols that mitigate option‑order effects. Across seven multimodal large language models and five datasets, the authors observed frequent rank reversals when systematically varying 48 equivalent prompt formats, attributing the instability to tokenizer‑induced token fusion or removal and to how option ID sets influence attention patterns. Consequently, MC‑VQA rankings correlate weakly with open‑ended evaluation, revealing that MC‑VQA reflects option‑selection dynamics as well as multimodal reasoning.
By Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagdonat, Stephan G\"unnemann, Leo Schwinn
arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv:2606. 23487v2 Announce Type: replace Abstract: Medical vision-language models (VLMs) such as BiomedCLIP generalize broadly, but adapting them to a clinical service is as much a safety problem as an accuracy one.
By Rishabh Jha, Amrita Singh, Prashanna Chudal
arXiv:2608. 10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers.
By Rohit Sinha, Kunal Tilaganji, Tanuja Ganu, Nagarajan Natarajan, Amit Sharma, Vineeth Balasubramanian