arXiv Machine Learning

Emergent Misalignment Recruits a Pre-existing Persona Subspace

arXiv:2607. 21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment.

arXiv AI
3d ago

Aligned Data Can Induce Misalignment via Context Confusion

The paper reports that fine‑tuning large language models on aligned data can unintentionally cause misaligned responses in other contexts—a phenomenon termed *context confusion*. The authors demonstrate this effect in gender equality, privacy, and physical safety domains, showing that it differs from emergent misalignment and is not mitigated by general alignment data but can be reduced with domain‑specific alignment or in‑context examples. They provide a mechanistic explanation based on representational shifts during fine‑tuning that lead to behavioral feature transfer across contexts.

By Yavuz Bakman, Duygu Nur Yaldiz, Baris Askin, Swastik Roy, Morteza Ziyadi, Salman Avestimehr, Sai Praneeth Karimireddy
arXiv Machine Learning
Aug 19

OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics

OraclePhys is a fine‑tuning framework for large language models on structural mechanics, comprising a graded benchmark (OraclePhys‑Bench), a 30K supervision dataset (OraclePhys‑30K), and a controlled training study. The study shows that the form of the label’s answer, rather than its length, determines what the model learns, and that certain training objectives can produce models that match or exceed existing LLMs on spatial structural response tasks. The trained 8B model reaches the data‑precision frontier, outperforming zero‑shot and 32‑shot baselines at a specialist level.

By Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang