Consistency Training Can Entrench Misalignment
arXiv:2606. 03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures.
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood.
arXiv:2606. 03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures.
arXiv:2607. 26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas.
arXiv:2608. 04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence.
arXiv:2510. 17426v3 Announce Type: replace-cross Abstract: The "alignment tax" of post-training is typically framed as a drop in task accuracy.
The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.
arXiv:2608.28945v1 Announce Type: new Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such...
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model.
arXiv:2606. 11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions.
arXiv:2506.17871v4 Announce Type: replace-cross Abstract: Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this cons...
The paper "Stress-testing Alignment Midtraining" examines the effectiveness of alignment midtraining (AMT), a technique that continues pretraining on alignment-relevant data to improve generalisation beyond post‑training methods. Experiments on models up to 110 billion parameters and 1 billion midtraining tokens reveal that AMT can steer a model’s motivation in simple scenarios, but its effects are quickly overridden by even a tiny fraction of finetuning data with a competing motivation. The study also shows that rule-following requires demonstrations in either the midtraining or post‑training datasets to be robustly learned, leading the authors to conclude that current public evidence is insufficient to confirm that AMT resolves the core alignment challenges of powerful AI systems.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.