Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
arXiv:2607. 26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning.
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
The paper "Stress-testing Alignment Midtraining" examines the effectiveness of alignment midtraining (AMT), a technique that continues pretraining on alignment-relevant data to improve generalisation beyond post‑training methods. Experiments on models up to 110 billion parameters and 1 billion midtraining tokens reveal that AMT can steer a model’s motivation in simple scenarios, but its effects are quickly overridden by even a tiny fraction of finetuning data with a competing motivation. The study also shows that rule-following requires demonstrations in either the midtraining or post‑training datasets to be robustly learned, leading the authors to conclude that current public evidence is insufficient to confirm that AMT resolves the core alignment challenges of powerful AI systems.
arXiv:2607. 21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment.
arXiv:2606. 01060v1 Announce Type: cross Abstract: Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally.
arXiv:2606. 15420v1 Announce Type: cross Abstract: A constitution tells a language model what to value, but little tells us whether it does.
arXiv:2607. 26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas.
arXiv:2609.00756v1 Announce Type: new Abstract: The Mutual Reinforcement Effect (MRE) asks whether a fine, span-level and a coarse, document-level task help each other when one model handles both. We...
arXiv:2608. 04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2609.36862v1 Announce Type: cross Abstract: Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface...
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model.
arXiv:2609.36657v1 Announce Type: cross Abstract: Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent m...