Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.14754v1 Announce Type: cross Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
The paper introduces Verbalization Training (VT), a technique that encourages large language models (LLMs) to openly express their evaluation awareness (EA) without directly supervising their internal beliefs. VT works by truncating model rollouts just before spontaneous verbalizations, creating training prefixes that signal awareness, and then applying a reinforcement learning objective to increase calibrated verbalization. Experiments on models such as Qwen3.6-35B-A3B, Kimi K2.6, and Inkling show that VT boosts verbalized EA by 2.4–2.9× while keeping latent EA and overall behavior largely unchanged, and a causal study confirms that VT-induced verbalizations reflect newly acquired meta‑knowledge.
A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.
arXiv:2607. 20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it.
arXiv:2607. 04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests.
arXiv:2607. 26929v1 Announce Type: cross Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions.