arXiv:2609.22206v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of multimodal tasks, yet understanding and quantify...
By Soroush Seifi, Vaggelis Dorovatas, Lin Li, Yarin Gal, Rahaf Aljundi
The paper argues that stochastic decoding, common in large language models, is not ideal for Visual Question Answering (VQA) because VQA is a closed‑ended task with head‑heavy answer distributions and epistemic uncertainty. The authors formalize how model calibration relates to predictive accuracy and identify conditions under which greedy decoding is optimal. Experiments across multiple benchmarks show greedy decoding outperforms stochastic sampling, and a new Greedy Decoding for Reasoning Models further improves multimodal reasoning performance.
By Boqi Chen, Xudong Liu, Yunke Ao, Jianing Qiu
Unified Multimodal Uncertain Inference (UMUI) is a new task that requires models to generate calibrated probability estimates for hypotheses conditioned on premises across text, audio, and video modalities. The authors create a human‑annotated evaluation set with scalar probability judgments for audio, visual, and audiovisual settings, and benchmark their approach on existing text and audio datasets. Their CLUE framework, which blends self‑consistent teacher calibration with distribution‑based confidence probing, enables a 3B‑parameter model to match or surpass zero‑shot baselines up to 32B parameters across all modalities.
By Dengjia Zhang, Alexander Martin, William Jurayj, Kenton Murray, Benjamin Van Durme, Reno Kriz
arXiv:2608.22232v1 Announce Type: new
Abstract: Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (ML...
By Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi
arXiv:2603. 28026v2 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs).
By Taeyun Roh, Suhyeong Park, Dongyoung Lee, Eunyeong Jo, Wonjune Jang, Junha Jung, Jaewoo Kang
arXiv:2605.27136v2 Announce Type: replace
Abstract: Uncertainty quantification (UQ) remains a critical challenge in Large Vision Language Models (LVLMs) for reliable predictions and real-world deploy...
By Joseph Hoche, David Brellmann, Gianni Franchi