arXiv AI By Hiskias Dingeto

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv Machine Learning
Aug 28

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

The paper investigates whether a language model’s own confidence can replace labeled data for teaching it to abstain from uncertain answers. By fine‑tuning models with LoRA to answer only when their frozen confidence is high and to say “I’m not sure” otherwise, the authors show that this label‑free approach matches label‑supervised abstention tuning on short‑form factual QA. The method works across six open‑weight models (1B‑8B) and is effective except for confidently wrong facts, which the confidence signal cannot flag.

By Ali Asaria, Tony Salomone, Deep Gandhi