AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,657 stories · RSS feed

arXiv AI
Jun 15

The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training

arXiv:2603. 10444v2 Announce Type: replace-cross Abstract: FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitudes, which inflate dynamic range and compress long-tail signals.

By Hengjie Cao, Zhendong Huang, Mengyi Chen, Yifeng Yang, Fang Dong, Anrui Chen, Ruijun Huang, Xin Zhang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Yixuan Chen, Li Shang
arXiv Machine Learning
Jun 15

A Low-Rank Subspace Analysis of LLM Interventions

arXiv:2606. 14388v1 Announce Type: new Abstract: Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors.

By Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu
arXiv AI
Jun 15

A Virtuous AI is an Existential Risk

arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.

By Guillermo Del Pinal, Youngchan Lee, Min Ohn
arXiv Machine Learning
Jun 15

Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

arXiv:2606. 14130v1 Announce Type: new Abstract: Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents.

By Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli
arXiv AI
Jun 15

Learning Developmental Scaffoldings to Guide Self-Organisation

arXiv:2605. 14998v3 Announce Type: replace Abstract: From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local interactions that collectively give rise to global structure without any blueprint of the outcome.

By Milton L. Montero, Elias Najarro, Jakob Schauser, Sebastian Risi
arXiv Machine Learning
Jun 15

Neural Variability Enhances Artificial Network Robustness

arXiv:2606. 13801v1 Announce Type: new Abstract: Neural responses in cortex exhibit substantial trial-to-trial variability in response to repeated stimuli, while peripheral sensory neurons respond far more consistently, leading many to wonder whether stochasticity may carry meaning.

By Robin Preble, Praveen Venkatesh, Stefan Mihalas, Kameron Decker Harris
arXiv AI
Jun 15

Recovering Stranded Discrimination in Knowledge Tracing: Per-Item Bias Correction via Empirical-Bayes Shrinkage

arXiv:2606. 14123v1 Announce Type: cross Abstract: Deployed knowledge-tracing models are typically frozen after training, yet systematic per-item logit bias arises, from limited per-item expressivity in backbone architectures and from post-deployment shifts in item properties, degrading prediction quality.

By Xiaoran Yan, Cheng Tang, Atsushi Shimada
arXiv Machine Learning
Jun 15

Generative Modeling of Bach-Style Symbolic Music: A Comparative Study of Autoregressive, Latent-Variable, and Adversarial Approaches

arXiv:2606. 13626v2 Announce Type: replace-cross Abstract: We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks.

By Dezhi Yu, Kyuil Lee, Yongkang Huang