AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv Machine Learning
Jun 2

Chaining 2-FWL GNNs for Combinatorial Graph Alignment

arXiv:2510. 03086v2 Announce Type: replace Abstract: For the combinatorial graph alignment problem (GAP) -- finding the node correspondence that maximizes the number of common edges (nce) between two unlabeled graphs -- properly initialized FAQ remains a strong classical baseline, while existing GNN approaches struggle in the purely structural setting.

By Marc Lelarge
arXiv Machine Learning
Jun 2

FedCF: Fair Federated Conformal Prediction

arXiv:2509. 22907v2 Announce Type: replace Abstract: Conformal Prediction (CP) is a widely used technique for quantifying uncertainty in machine learning models.

By Anutam Srinivasan, Aditya T. Vadlamani, Amin Meghrazi, Srinivasan Parthasarathy
arXiv AI
Jun 2

Universal Quantum Transformer

arXiv:2606. 00045v1 Announce Type: new Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact mathematical symmetries, such as modular arithmetic and non-commutative algebra.

By Sungyong Chung, Alireza Talebpour
arXiv Machine Learning
Jun 2

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

arXiv:2606. 01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals.

By Yuepeng Wang, Ken Kawano, Yongqi Zhou, Yoshihiko Fujisawa, Richard Weiss, Akifumi Wachi, Katsuki Fujisawa, Ying Chen, Mehrshad Sadria, Xin Liu, Kyoung-Sook Kim, Xiao Hu, Sebastien Gros, Xun Shen