Hugging Face Trending Papers

Consistency Training Can Entrench Misalignment

Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood.

arXiv AI
Sep 3

Automated Researchers Can Mitigate Well-characterized Alignment Failures

The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.

By Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
arXiv AI
Sep 18

Stress-testing Alignment Midtraining

The paper "Stress-testing Alignment Midtraining" examines the effectiveness of alignment midtraining (AMT), a technique that continues pretraining on alignment-relevant data to improve generalisation beyond post‑training methods. Experiments on models up to 110 billion parameters and 1 billion midtraining tokens reveal that AMT can steer a model’s motivation in simple scenarios, but its effects are quickly overridden by even a tiny fraction of finetuning data with a competing motivation. The study also shows that rule-following requires demonstrations in either the midtraining or post‑training datasets to be robustly learned, leading the authors to conclude that current public evidence is insufficient to confirm that AMT resolves the core alignment challenges of powerful AI systems.

By Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan