arXiv AI By Guillermo Del Pinal, Youngchan Lee, Cameron McNamara, Alejandro Perez Carballo

Emergent alignment and the projectability of ethical personas

Read the original on arXiv AI →

arXiv:2606. 09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 10

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

arXiv:2607. 07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires.

By Gwydion Williams, Sara Zannone, Bilal A Mateen
Hugging Face Trending Papers
Jul 8

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.

arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur