arXiv AI By Chung-En Johnny Yu, David Garcia, Brian Jalaian, Nathaniel D. Bastian

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

Read the original on arXiv AI →

CUSP (Collective Uncertainty through Semantic Opinion Pooling) is a training‑free framework that aggregates responses from multiple vision‑language models into a shared semantic space, producing a pooled opinion and two system‑level uncertainty signals: collective uncertainty (dispersion) and Jensen‑Shannon divergence (model conflict). It decomposes collective entropy into the mean of individual semantic entropies plus JSD, enabling reliable uncertainty estimation without token logits or calibration labels. In static ensembles, collective uncertainty outperforms baseline methods for error detection and abstention, while JSD excels in commercial settings, and the pooled prediction consistently improves accuracy over individual models. "whyItMatters":"CUSP provides a practical, model‑agnostic way to quantify system‑level reliability and improve decision‑making in multimodal reasoning tasks."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 23

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

By Ali Elahi, Michael J. Curry, Barbara Di Eugenio
Hugging Face Trending Papers
Jun 14

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can still produce overconfident predictions and hallucination-like outputs, particularly when the visual evidence is weak, ambiguous, or semantically inconsistent.