arXiv:2607. 21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment.
By Mohammed Suhail B Nadaf
TAME (Token Attribution and Masking for Emergent misalignment) is a three‑stage framework that identifies which training tokens drive harmful behavior in fine‑tuned language models. It first scores tokens by how much fine‑tuning increases their likelihood, then characterizes patterns among high‑attribution tokens, and finally validates them by masking during training. Experiments on Llama and Qwen show that masking the top‑attribution tokens reduces emergent misalignment by 23‑ to 36‑fold, while random masking has no effect.
By Md Rayhanul Masud, Md Rizwan Parvez
Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weig...
arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.
By Tung-Ling Li, Hongliang Liu, Yuhao Wu
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose
arXiv:2607. 10202v1 Announce Type: new Abstract: Cross-model comparisons read divergence in value dispositions as evidence that language models hold individuated values.
By Hong-In Won, Jinseok Jang, Hyoseop Kim
arXiv:2609.14759v1 Announce Type: cross
Abstract: Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model...
By Orion Reblitz-Richardson
The paper investigates PRM‑Pruned Fragment Grafting (PPFG), an inference‑time technique that extracts high‑reward prefixes from a pruned chain‑of‑thought and grafts them into a sibling decoding process. Experiments on Qwen2.5‑7B‑Instruct with Math‑Shepherd across 500 MATH problems and multiple seeds show that PPFG performs statistically indistinguishable from a parallel‑CoT baseline, with only 14% of grafts targeting genuinely struggling chains. The study extends across three language models, six benchmarks, and multiple PRM configurations, concluding that PPFG’s inertness is not due to heuristic specifics and providing an equivalence‑testing framework for mechanism nulls.
By Khawaja Murad ul Hassan, Mehran Ebrahimi
arXiv:2606. 17229v1 Announce Type: cross Abstract: A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone.
By Petr Nyoma
The study investigates whether a subliminal trait can persist across multiple generations of language‑model lineages. Three copies of Qwen2.5‑7B‑Instruct were trained for ten iterations, and the trait’s expression was measured via a keyword screen and an activation probe. Results show the trait remains detectable in all generations, though its behavioral expression diminishes, and it can exist internally without being overtly expressed when the system prompt is removed.
By Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen
arXiv:2607. 00006v1 Announce Type: cross Abstract: Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering.
By Shuaizhi Cheng