arXiv AI By Jiawei Zheng, Jiazhen Zhang

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

Read the original on arXiv AI →

arXiv:2607. 20529v1 Announce Type: cross Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

Reliable Fusion of Conflicting Experts

The paper introduces a probabilistic‑circuit framework for fusing opinions from multiple black‑box experts in noisy, conflict‑prone environments. It dynamically assigns context‑specific credibility to each expert, allowing reliable aggregation without needing access to their internal models or retraining. Experiments on multiple‑choice question answering with large language models show that this method outperforms individual models and static ensemble baselines, consistently improving predictive accuracy and decision reliability under disagreement.

By Pranuthi Tenali, Sahil Sidheekh, Saurabh Mathur, Vijayalakshmi Saravanan, Erik Blasch, Kristian Kersting, Sriraam Natarajan
arXiv AI
2d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano
arXiv AI
Jun 3

CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery

arXiv:2606. 03602v1 Announce Type: cross Abstract: Causal discovery from observational data remains challenging due to the fundamental limitations of purely statistical methods, such as statistical distinguishability within equivalence classes and sensitivity to finite sample sizes.

By Bo Peng, Kaiwen Wu, Sirui Chen, Zhiheng Wang, Yu Qiao, Chaochao Lu