AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv Machine Learning
Jun 2

The role of class encoding in neural collapse

arXiv:2606. 00344v1 Announce Type: new Abstract: Neural collapse is a structural property of the last-hidden-layer activations in neural network classification models, when trained beyond a zero classification error.

By Bastien Massion, Roy Makhlouf, Estelle Massart
arXiv AI
Jun 2

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation

arXiv:2606. 01897v1 Announce Type: new Abstract: Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC).

By Tianjiao Li, Kai Zhao, Xiang Li, Yang Liu, Huyang Sun
arXiv AI
Jun 2

Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training

arXiv:2601. 23220v2 Announce Type: replace-cross Abstract: Despite recent Multimodal Large Language Models (MLLMs)' linguistic prowess in medical diagnosis, we find even state-of-the-art MLLMs suffer from a critical perceptual deficit: geometric blindness.

By Anglin Liu, Ruichao Chen, Yi Lu, Hongxia Xu, Jintai Chen
arXiv AI
Jun 2

Prototype Transformer: Towards Language Model Architectures Interpretable by Design

arXiv:2602. 11852v2 Announce Type: replace Abstract: While state-of-the-art language models (LMs) surpass most humans in certain domains, their reasoning remains largely opaque, reducing trust and increasing the risk of deception and hallucination.

By Yordan Yordanov, Matteo Forasassi, Bayar Menzat, Ruizhi Wang, Chang Qi, Markus Kaltenberger, Amine M'Charrak, Tommaso Salvatori, Thomas Lukasiewicz
arXiv AI
Jun 2

How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning

arXiv:2606. 02119v1 Announce Type: cross Abstract: Machine unlearning aims to remove the influence of specific forget training data due to privacy, copyright or bias concerns while maintaining the model performance on the remaining retain data.

By Jiangwei Chen, Xinyuan Niu, Rachael Hwee Ling Sim, Zhengyuan Liu, Nancy F. Chen, Bryan Kian Hsiang Low
arXiv Machine Learning
Jun 2

DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole-Slide Image Survival Prediction

arXiv:2510. 00053v2 Announce Type: replace-cross Abstract: Pathology whole-slide images (WSIs) are widely used for cancer survival analysis because of their comprehensive histopathological information at both cellular and tissue levels, enabling quantitative, large-scale, and prognostically rich tumor feature analysis.

By Yucheng Xing, Ling Huang, Jingying Ma, Ruping Hong, Jiangdong Qiu, Pei Liu, Kai He, Huazhu Fu, Mengling Feng
arXiv Machine Learning
Jun 2

Near-Optimal Private Tests for Simple and MLR Hypotheses

arXiv:2601. 21959v2 Announce Type: replace-cross Abstract: We develop a near-optimal testing procedure under the framework of Gaussian differential privacy for simple as well as one- and two-sided tests under monotone likelihood ratio conditions.

By Yu-Wei Chen, Raghu Pasupathy, Jordan Awan