arXiv AI

Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication

arXiv:2605. 26397v2 Announce Type: replace-cross Abstract: Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication.

arXiv Computation and Language
Sep 23

From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health

The article surveys how large language models (LLMs) are being applied to mental health, outlining a three‑phase evolution: Phase I uses LLMs as passive information tools and pattern recognizers for assessment; Phase II employs them as empathetic conversationalists for stateless, in‑the‑moment interactions; Phase III aims to create longitudinal, personalized companions that act as stateful cognitive agents. It systematically reviews core technologies, agent architectures (Profile, Memory, Reasoning, Planning), and the datasets and benchmarks that support this progression, offering a coherent narrative and roadmap for future research. The survey also provides a curated resource list at https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.

By He Hu, Yucheng Zhou, Qianning Wang, Yingjian Zou, Chiyuan Ma, Juzheng Si, Jianzhuang Liu, Zitong Yu, Laizhong Cui, Fei Ma, Qi Tian
arXiv AI
Sep 18

Xeno-Interpretability: Investigating the Alien Minds of LLMs

The paper introduces the concept of xeno-interpretability, which studies internal distinctions in large language models that lack corresponding human concepts. It distinguishes between human‑interpretable and xeno‑semantic spaces, showing that LLMs possess a far larger internal representational space than can be captured by finite human descriptions. The authors propose an empirical program to identify and characterize these xeno‑representations, noting their potential to influence model behavior in ways that are not fully visible through human‑readable communication.

By F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu