AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,648 stories · RSS feed

arXiv Machine Learning
Jul 7

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

arXiv:2607. 02569v1 Announce Type: cross Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using RETFound, a self-supervised vision-transformer retinal foundation model used here as a frozen feature encoder, and the public APTOS 2019 and DDR diabetic retinopathy fundus image datasets.

By Karim Mardhani
arXiv AI
Jul 7

Privacy-Preserving Robustness Verification for Neural Networks

arXiv:2607. 05251v1 Announce Type: cross Abstract: Neural network verification and data privacy are inherently in tension: verification demands full access to model parameters and input data, yet both are increasingly restricted by privacy regulations and intellectual property constraints.

By Nianyun Song, Xiaokun Luan, Yu Guo, Rongfang Bie, Meng Sun, Xiyue Zhang
arXiv Machine Learning
Jul 7

Non-Asymptotic Error Bounds for SMC with Biased Proposals: Application to Conditional Diffusion Sampling

arXiv:2607. 04780v1 Announce Type: cross Abstract: Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow.

By Stanislas Strasman (SU, LPSM), Gabriel Victorino Cardoso (LPSM), Sylvain Le Corff (LPSM), Vincent Lemaire (LPSM), Antonio Ocello
arXiv Machine Learning
Jul 7

Functional Bilevel Optimization for Predictive Fairness

arXiv:2607. 05098v1 Announce Type: new Abstract: When sensitive attributes are continuous and high-dimensional $-$ demographic score vectors, posteriors over attributes, age or income profiles $-$ enforcing full statistical independence is often too restrictive, and existing relaxations rely on indirect dependence penalties or adversarial schemes that do not directly target the fairness-accuracy trade-off.

By Ieva Petrulionyte, Julien Mairal, Michael Arbel