A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
The paper demonstrates that a prompt’s influence is not inherent to the prompt itself but depends on the model, as prompts optimized for one model degrade on another and rankings shift under neutral reformatting. By examining a task‑free structural readout—specifically the fixed‑point behavior of a short‑window argmax map—the authors show that nine tokens of conditioning can move the fixed‑point fraction across most of its range, altering structural classes and model rankings, while instruction tuning has no effect. Attempts to explain this phenomenon through prefix length, content type, bidirectionality, or attention‑sink dominance all fail, indicating that the prompt‑model pair is the fundamental unit of explanation.
whyItMatters":"The study reveals that prompt effectiveness is model‑specific and that simple structural readouts can capture this interaction, challenging assumptions about prompt generality and guiding future prompt‑engineering efforts."
By Nicol\'as Vera Z\'u\~niga
arXiv:2607. 11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one.
By Bojie Li, Noah Shi
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
arXiv:2608. 04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models.
By Agatha Duzan, Asa Cooper Stickland
There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal.