Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that th...
The paper introduces Gated Activation Steering, an inference-time intervention that jointly mitigates hallucination and sycophancy in medical question answering. By learning separate steering directions from contrastive clinical pairs and applying them to specific attention heads, the method uses behavior‑specific gates to intervene only when needed. Experiments on EHR‑based clinical questions show that the 4‑billion‑parameter model with gated steering outperforms its unsteered counterpart and rivals larger models in resisting user pressure.
By Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi
arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.
By Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
arXiv:2609.32964v2 Announce Type: replace
Abstract: Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to...
By Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
The paper demonstrates that language models can acquire new capabilities from post‑training data even when the training text is unrelated to the target task. Using a method called Active Taskless Distillation (ATD), the authors show that a single word from a teacher model can transfer knowledge to a student model without any target‑task examples or teacher logits. Experiments on Qwen2.5-1.5B reveal significant performance gains on HumanEval+ and improvements in scientific knowledge, commonsense reasoning, and reading comprehension across various model families.
By Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong