arXiv AI By Volodymyr Ovcharov

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

Read the original on arXiv AI →

arXiv:2608. 15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
Sep 7

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

The paper investigates whether the object selected in a grounded language‑model pipeline actually reaches the reader, a failure that can break the handoff between stages. By auditing 600 HybridQA questions across three selector families, the authors find that exact key lookup and title matching recover the selected object in all 1,463 resolvable records, but body‑only BM25 omits it in 26.6% of cases at cutoff five, while hybrid retrieval with reranking omits it only 1.0%. The study also shows that misalignment between selected and retrieved objects can reduce exact match scores by up to 31 points, and introduces the Returned‑Object Profile (ROP) as a tool for reproducible auditing.

By Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu
arXiv AI
2d ago

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

The paper demonstrates that reading a large language model (LLM) judge’s verdict from the logits of its first generated token—an approach used in constrained decoding and likelihood‑scoring evaluation—introduces a significant distortion in position bias. Because judges do not always start with a verdict token (12–49% of cases for Qwen3 judges and <3% for Llama‑3.1‑8B and Phi‑3.5‑mini), this readout often returns the first response rather than a true judgment, inflating position bias by up to 42 points while barely affecting judge accuracy. The authors recommend reporting the frequency with which a judge leads with a verdict token to provide a more accurate assessment of position bias.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli