arXiv:2609.39848v1 Announce Type: new
Abstract: Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundatio...
By Mingyue Ma, Zongbo Han, Changqing Zhang, Guangyu Wang
arXiv:2603. 03989v2 Announce Type: replace-cross Abstract: When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns.
By Qianpu Chen, Derya Soydaner, Rob Saunders
arXiv:2609.15180v1 Announce Type: new
Abstract: Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliabl...
By Mingcheng Zhu, Jinning Liang, Tingting Zhu
arXiv:2606. 15767v1 Announce Type: cross Abstract: Understanding when and why deep neural networks are uncertain is crucial for deploying reliable machine learning systems in safety-critical domains.
By Dong Hyun Jeong, Feng Chen, Jin-Hee Cho, Lance M. Kaplan, Audun J{\o}sang, Soo-Yeon Ji
arXiv:2601. 21944v3 Announce Type: replace Abstract: The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability.
By Konstantinos P. Panousis, Diego Marcos
arXiv:2608.30789v1 Announce Type: new
Abstract: Supervised deep learning methods enable the rapid processing of ecological image data, but depend on a costly annotation process. Consequently, trainin...
By Leonard Hockerts, Peter S. Stewart, Sarthak Arora, Tiffany J. Vlaar
arXiv:2604. 25077v2 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots.
By Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh
Video foundation models now match human accuracy on physical‑reasoning benchmarks, but a new distributional evaluation framework shows that their predictions diverge markedly from human judgments. On the Physion benchmark, ViT‑L models such as V‑JEPA2, VideoMAE‑v2, and DINOv2 achieve near‑human accuracy yet exhibit a 26.4% model‑human disagreement, far above the 4.8% human‑human disagreement, and lower agreement (kappa ~0.48 vs. 0.91). The divergence varies by task: models excel at geometric reasoning but lag on gravitational dynamics and causal chains, indicating they rely on statistical regularities rather than explicit forward simulation.
By Fanhong Li, Shurui Zheng, Zi Yin, Junbo Cui, Lei Ji, Jia Liu
arXiv:2606. 05376v1 Announce Type: new Abstract: Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challenging disagreements across human annotators.
By Jingyao Wu, Ashley Wang, Keane Ong, Paul Pu Liang, Rosalind Picard
arXiv:2511. 14117v2 Announce Type: replace Abstract: Supervised classifiers output a distribution over classes but are typically trained against a single label obtained by collapsing multiple annotators into a majority vote.
By Agamdeep Singh, Ashish Tiwari, Hosein Hasanbeig, Priyanshu Gupta
arXiv:2607. 22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation.
By Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao