AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,365 stories · RSS feed

arXiv Machine Learning
Jun 8

Model Recycling Framework for Multi-Source Data-Free Supervised Transfer Learning

arXiv:2508. 02039v2 Announce Type: replace Abstract: Increasing concerns for data privacy and other difficulties associated with retrieving source data for model training have created the need for source-free transfer learning, in which one only has access to pre-trained models instead of data from the original source domains.

By Sijia Wang, Ricardo Henao
arXiv Machine Learning
Jun 8

Automatic, Debiased, and Invariant Counterfactual Generation under General Interventions

arXiv:2606. 07399v1 Announce Type: cross Abstract: Generative models for counterfactual outcomes have great potential to support decision-making under complex interventions, but existing approaches are limited by unstable estimation, poor generalization across environments, and bias from nuisance model misspecification.

By Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina
arXiv Machine Learning
Jun 8

Federated Foundation Models over Vehicular Networks

arXiv:2606. 06786v1 Announce Type: new Abstract: This paper presents a forward-looking vision for integrating the emerging multi-modal multi-task federated foundation models (M3T FedFMs) into vehicular networks, with the goal of unifying the expressive power of multi-modal multi-task foundation models (M3T FMs) with the privacy-preserving and distributed learning capabilities of federated learning (FL).

By Kasra Borazjani, Fardis Nadimi, Payam Abdisarabshali, Owen Palinski, Allan Salihovic, Dinh Nguyen, Minghui Liwang, Seyyedali Hosseinalipour
arXiv AI
Jun 8

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

arXiv:2602. 02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and initial evidence of improved jailbreak robustness.

By Eliron Rahimi, Elad Hirshel, Rom Himelstein, Amit LeVi, Avi Mendelson, Chaim Baskin
arXiv AI
Jun 8

FIGMA: Towards FIne-Grained Music retrievAl

arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.

By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
arXiv Machine Learning
Jun 8

Covariance Shrinkage via Stochastic Interpolation

arXiv:2606. 07382v1 Announce Type: new Abstract: We recast classical shrinkage of high-dimensional covariance estimators as empirical risk minimization over a parametric stochastic interpolant between a source and a target distribution.

By Mathieu Chalvidal, Florentin Coeurdoux, Eric Vanden-Eijnden
arXiv Machine Learning
Jun 8

Deep Single-Index Fr\'echet Regression

arXiv:2606. 06957v1 Announce Type: cross Abstract: Predicting outputs that are located in non-Euclidean spaces, such as probability distributions, networks, and symmetric positive-definite matrices, is becoming increasingly important in modern data analysis, particularly when inputs are high-dimensional.

By Muqing Cui, Yidong Zhou, Su I Iao, Hans-Georg M\"uller