arXiv:2608.20442v1 Announce Type: new
Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically...
By Qinyang Xu
arXiv:2606. 00831v1 Announce Type: new Abstract: Subliminal learning is a phenomenon where language models can transmit behavioral traits to other models through seemingly innocuous data (Cloud et al.
By Todd Nief, Harvey Yiyun Fu, Mark Muchane, Ari Holtzman
arXiv:2609.22215v1 Announce Type: cross
Abstract: Knowledge distillation can transmit unintended behavioral traits from a teacher model to a student through training data that appear semantically unr...
By Atsushi Yanagisawa, Brendan Gho, Rajendran Ramesh Babu Manoj Narender, Kevin Zhu, Madhur Panwar, Antonio Mari
The asymmetry between language production and perception has been well-documented in psycholinguistics. Whether large language models (LLMs) exhibit a functionally analogous distinction remains an open question, particularly given that LLMs rely on the same underlying mechanism (next-token prediction) for both input and output processing.
arXiv:2606. 00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.
By Camila Blank, Agam Bhatia, Senthooran Rajamanoharan, Arthur Conmy, Neel Nanda
arXiv:2606. 24952v1 Announce Type: cross Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should be able to modify it.
By Cosimo Galeone, Anna Ettorre, Minsu Park, Giuseppe Ettorre, Daniele Ligorio
arXiv:2608. 16085v1 Announce Type: cross Abstract: Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state.
By Simon Lam-Muir
arXiv:2603.22161v3 Announce Type: replace
Abstract: Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstr...
By Dharshan Kumaran, Nathaniel Daw, Simon Osindero, Petar Veli\v{c}kovi\'c, Viorica Patraucean
The paper investigates subliminal learning, where hidden traits from a teacher model are transferred to a student during distillation. It introduces trait‑direction drift as the underlying mechanism, showing that biased generation creates measurable preference gaps that accumulate into behavioral transfer during fine‑tuning. The authors propose probe‑space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction, significantly reducing hidden‑trait transfer while maintaining task performance.
By Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang
The paper demonstrates that a prompt’s influence is not inherent to the prompt itself but depends on the model, as prompts optimized for one model degrade on another and rankings shift under neutral reformatting. By examining a task‑free structural readout—specifically the fixed‑point behavior of a short‑window argmax map—the authors show that nine tokens of conditioning can move the fixed‑point fraction across most of its range, altering structural classes and model rankings, while instruction tuning has no effect. Attempts to explain this phenomenon through prefix length, content type, bidirectionality, or attention‑sink dominance all fail, indicating that the prompt‑model pair is the fundamental unit of explanation.
whyItMatters":"The study reveals that prompt effectiveness is model‑specific and that simple structural readouts can capture this interaction, challenging assumptions about prompt generality and guiding future prompt‑engineering efforts."
By Nicol\'as Vera Z\'u\~niga
arXiv:2608. 12935v1 Announce Type: new Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means.
By Lei You
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil