The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.
By Atul Anand
arXiv:2609. 22880v1 Announce Type: cross Abstract: LLM rerankers add of the order of \$0.
By Andre Bacellar
arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.
By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
The paper investigates whether the object selected in a grounded language‑model pipeline actually reaches the reader, a failure that can break the handoff between stages. By auditing 600 HybridQA questions across three selector families, the authors find that exact key lookup and title matching recover the selected object in all 1,463 resolvable records, but body‑only BM25 omits it in 26.6% of cases at cutoff five, while hybrid retrieval with reranking omits it only 1.0%. The study also shows that misalignment between selected and retrieved objects can reduce exact match scores by up to 31 points, and introduces the Returned‑Object Profile (ROP) as a tool for reproducible auditing.
By Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu
The paper demonstrates that reading a large language model (LLM) judge’s verdict from the logits of its first generated token—an approach used in constrained decoding and likelihood‑scoring evaluation—introduces a significant distortion in position bias. Because judges do not always start with a verdict token (12–49% of cases for Qwen3 judges and <3% for Llama‑3.1‑8B and Phi‑3.5‑mini), this readout often returns the first response rather than a true judgment, inflating position bias by up to 42 points while barely affecting judge accuracy. The authors recommend reporting the frequency with which a judge leads with a verdict token to provide a more accurate assessment of position bias.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv:2609.37993v1 Announce Type: cross
Abstract: The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence...
By Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch
The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.
By Alden Do Rosario, Hussein Younes, Felipe Pires
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.
By Idil Gozel
arXiv:2607. 02104v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.
By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question.
"whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."
By Peter Potash