AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,365 stories · RSS feed

arXiv Machine Learning
Jun 8

Explaining Unsupervised Disease Staging in Huntington's Disease: Insights into Model Representations and Clusters

arXiv:2606. 07135v1 Announce Type: new Abstract: Huntington's disease (HD) is a progressive neurodegenerative disorder that affects motor, cognitive, and behavioral functions, where accurate characterization of disease progression remains essential to improve patient outcome and quality of life.

By Lubna Mahmoud Abu Zohair, Hind Zantout
arXiv AI
Jun 8

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

arXiv:2606. 07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded in a human-driven vehicle paradigm.

By Chaitanya Shinde, Hadi Hajieghrary, Paul Schmitt, Adam Shoemaker, Bodo Seifert, Steve Kenner
arXiv Machine Learning
Jun 8

Learning Fair Demand Models

arXiv:2606. 06830v1 Announce Type: cross Abstract: Data-driven pricing is increasingly prevalent in sectors such as airlines, lending, insurance, and retail.

By Adam N. Elmachtoub, Hyemi Kim, Jonathan Y. Tan
arXiv AI
Jun 8

FIGMA: Towards FIne-Grained Music retrievAl

arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.

By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami