AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,801 stories · RSS feed

arXiv AI
Jun 10

MoE Enhanced Federated Learning for Spatiotemporal Prediction

arXiv:2606. 10499v1 Announce Type: cross Abstract: Traffic prediction is fundamental to intelligent transportation systems and urban computing, yet many cities continue to suffer from traffic data scarcity due to limited sensor deployment and uneven urban development.

By Zhehao Dai, Xiao Han, Zhaolin Deng, Zijian Zhang, Xiangyu Zhao, Guojiang Shen, Xiangjie Kong
arXiv AI
Jun 10

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv:2606. 10569v1 Announce Type: cross Abstract: Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target.

By Dorcas Chia Ern Chua, Karen Myn Hui Lee, Jia Yue Tan, Zhen Xue Gue, Norzalena Abdul Hamid, Azima Binti Azmi, Keat Mei Yeong, Aizat Izyani binti Mujab, Hafsah Noor Azam, Chee Guo Khoo, Han Ying Lim, Chee Seng Chan
arXiv Machine Learning
Jun 10

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.

By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv Machine Learning
Jun 10

Interpretable deep convolutional model for nonlinear multivariate time series in complex systems

arXiv:2501. 04339v2 Announce Type: replace-cross Abstract: We introduce the Deep Convolutional Interpreter for Time Series (DCIts), a deep-learning architecture for nonlinear multivariate time series that provides sample-specific, locally interpretable descriptions of the underlying interaction structure.

By Domjan Baric, Davor Horvatic