arXiv AI

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

CUSP (Collective Uncertainty through Semantic Opinion Pooling) is a training‑free framework that aggregates responses from multiple vision‑language models into a shared semantic space, producing a pooled opinion and two system‑level uncertainty signals: collective uncertainty (dispersion) and Jensen‑Shannon divergence (model conflict). It decomposes collective entropy into the mean of individual semantic entropies plus JSD, enabling reliable uncertainty estimation without token logits or calibration labels. In static ensembles, collective uncertainty outperforms baseline methods for error detection and abstention, while JSD excels in commercial settings, and the pooled prediction consistently improves accuracy over individual models. "whyItMatters":"CUSP provides a practical, model‑agnostic way to quantify system‑level reliability and improve decision‑making in multimodal reasoning tasks."

arXiv Machine Learning
Sep 23

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

By Ali Elahi, Michael J. Curry, Barbara Di Eugenio
Hugging Face Trending Papers
Jun 14

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can still produce overconfident predictions and hallucination-like outputs, particularly when the visual evidence is weak, ambiguous, or semantically inconsistent.

Hugging Face Trending Papers
5d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.