arXiv AI By Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Read the original on arXiv AI →

arXiv:2607. 24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano
arXiv AI
3d ago

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.

By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan