AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,776 stories · RSS feed

arXiv Machine Learning
Aug 4

Latent-Regime Bias Auditing for Volatility Forecasting

arXiv:2608. 01599v1 Announce Type: new Abstract: Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management.

By Arthur Chagas, Pedro Bento, Yan Aquino, Arthur Buzelin, Wagner Meira Jr., Cristiano Arbex Valle
arXiv Machine Learning
Aug 4

Neural operator learning for collision-aware trajectory planning of spacecraft swarms

arXiv:2608. 00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and learning-based planners rarely transfer across swarm sizes or debris densities.

By Sidhdharth D. Sikka, Suyi Gao, Zehui Lu, Rongjie Lai, Shaoshuai Mou
arXiv Machine Learning
Aug 4

Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning

arXiv:2608. 02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization.

By Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
arXiv Machine Learning
Aug 4

Behavioural Analysis of Alignment Faking

arXiv:2605. 27681v2 Announce Type: replace-cross Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences.

By Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King, Alan Cooney
arXiv Machine Learning
Aug 4

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

arXiv:2603. 07475v4 Announce Type: replace-cross Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are trained through full-sequence denoising.

By Raghavv Goel, Risheek Garrepalli, Sudhanshu Agrawal, Chris Lott, Mingu Lee, Fatih Porikli
arXiv Machine Learning
Aug 4

Meritocratic Fairness via $K$-Shapley Values in Budgeted Combinatorial Bandits with Full-Bandit Feedback

arXiv:2605. 00762v2 Announce Type: replace Abstract: We study meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback, where a learner selects at most $K$ arms per time step and observes only the noisy aggregate reward of the selected set.

By Shradha Sharma, Shweta Jain, Swapnil Dhamal