Predicting Multi-View Rashomon Representation: Can We Learn Where Models Disagree?
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.39848v1 Announce Type: new Abstract: Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundatio...
arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...
The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.
arXiv:2507. 06722v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations.
The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.
The paper introduces a new consistency criterion for auditing decision systems that combines ensemble margin with local prediction variability to address predictive multiplicity, or the Rashomon effect. It shows that finite ensembles converge to the expected model’s consistency score as ensemble size and sample count grow, and demonstrates that ensembling models from the Rashomon set reduces unchecked incorrect predictions while keeping diversions moderate. Experiments on transformer and fine‑tuned language models for NLP and tabular classification confirm the method’s effectiveness and stronger alignment with existing multiplicity metrics.