AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv Machine Learning
Jun 2

How Neural Losses Shape VAE Latents

arXiv:2606. 00635v1 Announce Type: new Abstract: Modern VAEs are rarely trained with the pointwise likelihood implied by the standard $\beta$-VAE objective.

By Giorgio Strano, Luca Cerovaz, Michele Mancusi, Tommaso Mencattini, Emanuele Rodol\`a
arXiv AI
Jun 2

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation

arXiv:2606. 01897v1 Announce Type: new Abstract: Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC).

By Tianjiao Li, Kai Zhao, Xiang Li, Yang Liu, Huyang Sun
arXiv AI
Jun 2

UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment

arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.

By Zheng Wang, Shuo Wang, Junhong Wang
arXiv Machine Learning
Jun 2

Near-Optimal Private Tests for Simple and MLR Hypotheses

arXiv:2601. 21959v2 Announce Type: replace-cross Abstract: We develop a near-optimal testing procedure under the framework of Gaussian differential privacy for simple as well as one- and two-sided tests under monotone likelihood ratio conditions.

By Yu-Wei Chen, Raghu Pasupathy, Jordan Awan
arXiv AI
Jun 2

A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models

arXiv:2601. 17952v2 Announce Type: replace-cross Abstract: Interpretability remains a key challenge for deploying language models (LM) in clinical settings such as progression diagnosis of Alzheimer disease, where early and trustworthy predictions are essential.

By Michail Mamalakis, Tiago Azevedo, Cristian Cosentino, Chiara D'Ercoli, Subati Abulikemu, Zhongtian Sun, Richard Bethlehem, Pietro Lio
arXiv AI
Jun 2

Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment

arXiv:2606. 02322v1 Announce Type: cross Abstract: In dynamic environments, large language models need to keep adapting to new tasks, but continual learning often suffers from forgetting, limited transfer, and vulnerability to adversarial perturbations.

By Ran Liu, Min Yu, Mingqi Liu, Jianguo Jiang, Gang Li, Rongsheng Li, Ning Li, Zhen Xu, Weiqing Huang, Ming Liu