arXiv AI By Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv AI
Aug 18

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

The article argues that current evaluations of large language models’ moral competence focus mainly on whether outputs align with human moral values—the so‑called moral value problem—while neglecting the moral norm problem, which concerns the models’ ability to identify and apply context‑sensitive moral norms. It attributes this imbalance to the field’s reliance on descriptive ethics frameworks that emphasize value representation over normative application. The authors review existing benchmarks, highlight three gaps—lack of ground‑truth norm data, insufficient evaluation of intermediate reasoning, and limited focus on context‑relevant features—and propose a research agenda to develop formal normative representations, expert‑annotated datasets, and evaluation protocols that distinguish between value‑level and norm‑level competence.

By Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh