AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv AI
Jun 4

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

arXiv:2606. 04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep RL may demonstrate susceptibility to transition perturbations that result in unknown or unsafe behaviour.

By Mohit Prashant, Arvind Easwaran
arXiv AI
Jun 4

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

arXiv:2605. 07724v2 Announce Type: replace-cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective.

By Ali Falahati, Mohammad Mohammadi Amiri, Kate Larson, Lukasz Golab