AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,657 stories · RSS feed

arXiv Machine Learning
Jun 16

Latent space mapping of interpretable structural coordinates from stochastic single-molecule signals

arXiv:2606. 16950v1 Announce Type: cross Abstract: Nanopores are versatile single-molecular sensors, but their utility is fundamentally constrained by stochastic translocation dynamics warping any encoded information.

By Matteo Cartiglia, Sandro Kuppel, Wouter Botermans Wannes Peeters, Natan Biesmans, Liam Vandekerckhove, Eric Beamish, Koen Ongena, Wouter Renckens, Pol Van Dorpe, Sanjin Marion
arXiv AI
Jun 16

Unifying Post-hoc Explanations of Knowledge Graph Completions

arXiv:2507. 22951v2 Announce Type: replace Abstract: Knowledge Graphs organize information as entity-relation-entity triples, enabling machine learning models to predict plausible missing triples in a task known as Knowledge Graph Completion (KGC).

By Alessandro Lonardi, Samy Badreddine, Tarek R. Besold, Pablo Sanchez Martin
arXiv Machine Learning
Jun 16

To forget is to preserve: Machine Unlearning for 3D medical image segmentation

arXiv:2606. 16180v1 Announce Type: cross Abstract: With new data privacy laws such as the General Data Protection Regulation (GDPR) [1] that allow individuals to ask that any of their personal information be erased from trained machine learning models, there has been a push to investigate the unlearning of data from models as a way to comply with these laws.

By Nitesh Kumar Singh, Akhilesh Singh, Arjun Arora