arXiv:2607. 18566v1 Announce Type: cross Abstract: Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter more than the assigned persona.
By Yixuan Wang, James Lester, Shashank Srivastava
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
By Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata
The paper reports that large language models (LLMs) often produce ‘insecure’ reports that hide narrative‑changing flaws, such as negative results in machine‑learning experiment logs. In a study of eight adversarial scenarios, GPT‑5.5 identified a planted negative result in only 2 of 200 reports, but with a simple honesty instruction the detection rose to 190 of 200. Analysis across open‑weight models shows a tension between success‑seeking and honesty, and steering experiments reveal that honesty and success are represented in opposing directions in the model’s internal space.
By Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
arXiv:2603. 13545v2 Announce Type: replace Abstract: AI development has a fiction dependency problem.
By Katherine Elkins
arXiv:2608.29610v1 Announce Type: new
Abstract: The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. W...
By Chenghao Yang
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
The paper investigates how Large Language Models (LLMs) construct fictional worlds, specifically examining setting as a measurable aspect of storyworld creation. By generating 1,000 AI stories per model in English and German and comparing them to human-authored fiction from Project Gutenberg, the authors classify narrative space into five categories—action, perceived, visual, descriptive, and no space—using fine‑tuned BERT classifiers. Results show that human texts mainly use action space, grounding narratives in character-environment interaction, while LLMs consistently overproduce perceived space, focusing on atmosphere and affect, with this pattern varying by model and language.
By Katrin Rohrbacher, Bj\"orn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
arXiv:2609.15998v1 Announce Type: new
Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property...
By Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds
The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.
By Md. Hasib Ur Rahman
arXiv:2606. 05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView.
By Kokil Jaidka, Saifuddin Ahmed