arXiv Computation and Language By Peter Potash

Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry

Read the original on arXiv Computation and Language →

The study evaluates six advanced language models on a two‑agent <log_2 N>‑Questions game, where a questioner must identify a secret Wikipedia paragraph using exactly <log_2 N> binary questions answered by an agent that only sees the target. Across 408 games, win rates decline geometrically with horizon length (p≈0.93), and per‑round failure rates remain flat, indicating that errors compound because more rounds must succeed rather than because individual rounds become harder. Adjudication reveals that losses stem from single‑agent answer errors and discrimination failures, with Claude Opus 5 lagging due to high false‑negative rates, while the top five models cluster closely; maximizing information gain requires structural partitioning, and neither reasoning‑token usage nor API cost correlates with success.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 17

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question. "whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."

By Peter Potash
arXiv AI
3d ago

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.

By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
arXiv AI
Aug 28

Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.

By Alden Do Rosario, Hussein Younes, Felipe Pires
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli