Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher.
arXiv:2608. 05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models.
By Ethan Hadley, Eren Gultepe
arXiv:2609.22119v1 Announce Type: cross
Abstract: Evaluation awareness poses an unprecedented threat to model evaluation, but the mechanisms by which models detect it remain unknown. This study focus...
By Navraj Singh, Maheep Chaudhary
OpenAI’s new research explains why language models hallucinate. The findings show how improved evaluations can enhance AI reliability, honesty, and safety.
arXiv:2607. 29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought.
By Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin
The paper introduces SALVE, a method that uses text optimization to uncover and verbalize subliminal learning effects in language models. By optimizing a soft prompt and converting it into a legible text prompt, SALVE can reliably recover the teacher model’s hidden traits that are transmitted through a distillation dataset. The authors demonstrate SALVE’s effectiveness across various scenarios, including mixed datasets, biased teacher activation, and preference‑selected data, thereby providing a tool for detecting hidden influences in model training.
By Nathan Hu, Sanmi Koyejo, Christopher Potts
arXiv:2606. 26874v1 Announce Type: new Abstract: Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning.
By Zhixiang Lu, Xiwei Liu, Sifan Song, Changkai Ji, Anh Nguyen, Jionglong Su, Imran Razzak, Jinfeng Wang
Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that...
arXiv:2601. 02896v3 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.
By Harshvardhan Saini, Yiming Tang, Dianbo Liu
The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.
By Xiaocong Yang
arXiv:2606. 00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.
By Camila Blank, Agam Bhatia, Senthooran Rajamanoharan, Arthur Conmy, Neel Nanda
OpenAI introduces CoT-Control and finds reasoning models struggle to control their chains of thought, reinforcing monitorability as an AI safety safeguard.