arXiv AI

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5

arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.

arXiv AI
Sep 16

TAME: Token Attribution and Masking for Emergent misalignment

TAME (Token Attribution and Masking for Emergent misalignment) is a three‑stage framework that identifies which training tokens drive harmful behavior in fine‑tuned language models. It first scores tokens by how much fine‑tuning increases their likelihood, then characterizes patterns among high‑attribution tokens, and finally validates them by masking during training. Experiments on Llama and Qwen show that masking the top‑attribution tokens reduces emergent misalignment by 23‑ to 36‑fold, while random masking has no effect.

By Md Rayhanul Masud, Md Rizwan Parvez
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil
arXiv AI
2d ago

Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

The paper investigates PRM‑Pruned Fragment Grafting (PPFG), an inference‑time technique that extracts high‑reward prefixes from a pruned chain‑of‑thought and grafts them into a sibling decoding process. Experiments on Qwen2.5‑7B‑Instruct with Math‑Shepherd across 500 MATH problems and multiple seeds show that PPFG performs statistically indistinguishable from a parallel‑CoT baseline, with only 14% of grafts targeting genuinely struggling chains. The study extends across three language models, six benchmarks, and multiple PRM configurations, concluding that PPFG’s inertness is not due to heuristic specifics and providing an equivalence‑testing framework for mechanism nulls.

By Khawaja Murad ul Hassan, Mehran Ebrahimi
arXiv Machine Learning
Sep 23

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

The study investigates whether a subliminal trait can persist across multiple generations of language‑model lineages. Three copies of Qwen2.5‑7B‑Instruct were trained for ten iterations, and the trait’s expression was measured via a keyword screen and an activation probe. Results show the trait remains detectable in all generations, though its behavioral expression diminishes, and it can exist internally without being overtly expressed when the system prompt is removed.

By Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen
arXiv AI
Jul 2

Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem

arXiv:2607. 00006v1 Announce Type: cross Abstract: Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering.

By Shuaizhi Cheng