When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
arXiv:2608. 08029v1 Announce Type: cross Abstract: Khatri et al.
arXiv:2609.22478v1 Announce Type: cross Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We...
arXiv:2608. 11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance.
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
arXiv:2609.35815v1 Announce Type: cross Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated co...
arXiv:2606. 13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized.
Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across five aggregation methods, five datasets, five architectures, and four recorded conditions: clean, sign-flipping, Gaussian, and BadNets.
arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.
arXiv:2609.07944v1 Announce Type: new Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recove...