AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv Machine Learning
Jun 3

Correcting Neural Operator Spectral Bias via Diffusion Posterior Sampling with Sparse Observations

arXiv:2606. 03936v1 Announce Type: new Abstract: Neural operator surrogates (NO) approximate PDE solutions orders of magnitude faster than numerical solvers, but suffer from spectral bias: high-frequency content is systematically attenuated, limiting reliability where fine-scale structure matters.

By Niccol\`o Perrone, Fanny Lehmann, Stefania Fresca, Filippo Gatti
arXiv AI
Jun 3

RobotValues: Evaluating Household Robots When Human Values Conflict

arXiv:2606. 03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robots are expected to choose actions that prioritize other values than task success, such as human autonomy, efficiency, or social appropriateness.

By Jongwook Han, Hyeongjin Kim, Yohan Jo
arXiv AI
Jun 3

PURGE: Projected Unlearning via Retain-Guided Erasure

arXiv:2606. 03808v1 Announce Type: cross Abstract: We propose PURGE, a machine unlearning algorithm built on a simple but an under-exploited observation: continual learning (CL) and machine unlearning (MU) which are fundamentally dual problems.

By Vedant Jawandhia, Daksh Ahuja, Ghufran Alam Siddiqui, Prashant Trivedi, Yash Sinha, Pratik Narang