When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
arXiv:2606. 27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy.
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
arXiv:2608. 14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt.
arXiv:2608. 14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost.
The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.
arXiv:2609.37647v1 Announce Type: cross Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
arXiv:2608.23023v2 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall...
arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.
The paper investigates why large‑language‑model (LLM) routers—systems that select the best model for each query—often fail to outperform a single best model. By evaluating 14 models on 294 questions across seven task types and three languages, the authors find that a simple static mapping of task type to model improves 21 of the 29 questions that routing could potentially solve, and that learned routers do not significantly exceed this performance. The study highlights that most routing gains stem from task‑type specialization rather than complex learned decision rules.
arXiv:2608. 16391v1 Announce Type: cross Abstract: As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem.
arXiv:2609.39229v1 Announce Type: cross Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question. "whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."
arXiv:2608. 12895v1 Announce Type: new Abstract: Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested.