AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,811 stories · RSS feed

arXiv AI
Jun 8

Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation

arXiv:2606. 06869v1 Announce Type: new Abstract: Aim: Existing AI-assisted traditional Chinese medicine diagnostic tools suffer from opaque reasoning processes, passive interaction, and limited treatment plan presentation.

By Yunhan Wang, Yuda Wang, Zhiying Tu, Mingqiang Song, Li Song, Kun Li, Dianhui Chu, Bolin Zhang
arXiv Machine Learning
Jun 8

Model Recycling Framework for Multi-Source Data-Free Supervised Transfer Learning

arXiv:2508. 02039v2 Announce Type: replace Abstract: Increasing concerns for data privacy and other difficulties associated with retrieving source data for model training have created the need for source-free transfer learning, in which one only has access to pre-trained models instead of data from the original source domains.

By Sijia Wang, Ricardo Henao
arXiv Machine Learning
Jun 8

Learning Fair Demand Models

arXiv:2606. 06830v1 Announce Type: cross Abstract: Data-driven pricing is increasingly prevalent in sectors such as airlines, lending, insurance, and retail.

By Adam N. Elmachtoub, Hyemi Kim, Jonathan Y. Tan
arXiv Machine Learning
Jun 8

Automatic, Debiased, and Invariant Counterfactual Generation under General Interventions

arXiv:2606. 07399v1 Announce Type: cross Abstract: Generative models for counterfactual outcomes have great potential to support decision-making under complex interventions, but existing approaches are limited by unstable estimation, poor generalization across environments, and bias from nuisance model misspecification.

By Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina