AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,195 stories · RSS feed

arXiv Machine Learning
Jul 24

Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry

arXiv:2607. 20778v1 Announce Type: new Abstract: Weather forecasting foundation models (FMs) are increasingly fine-tuned to predict air quality, offering fast global pollution forecasts at lower computational cost than conventional chemical transport models.

By Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm, Makoto M. Kelp
arXiv Machine Learning
Jul 24

From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python

arXiv:2607. 21069v1 Announce Type: new Abstract: The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal.

By Muntasir Adnan, Manile Srun, Carlos C. N. Kuhn
arXiv Machine Learning
Jul 24

HERMES: Heterogeneous Edge-Relational Multi-Head Embedded SSM Attention for Traffic Conflict Prediction at Signalized Intersections

arXiv:2607. 20505v1 Announce Type: cross Abstract: Surrogate safety measures (SSMs) enable proactive traffic safety assessment, but many existing methods evaluate pairwise interactions independently or flatten multi-agent scenes into fixed feature vectors, limiting their ability to represent heterogeneous interaction structure and evolving scene-level risk.

By Md Monzurul Islam, Subasish Das
arXiv Machine Learning
Jul 24

SPECTRA: State-Space Exogenous Context and Temporal-Frequency Resolution Architecture for Probabilistic Energy Forecasting

arXiv:2607. 20587v1 Announce Type: cross Abstract: Modern power systems increasingly require probabilistic forecasts amid interacting uncertainties from renewable intermittency, flexible demand, market volatility, and weather-dependent generation.

By Hang Ye, Xinyan Jiang, Yuedong Shi, Yangxin Zhu, Jianming Wei, Tian Zheng, Xiaoying Zheng, Yongxin Zhu
arXiv Machine Learning
Jul 24

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

arXiv:2607. 21137v1 Announce Type: cross Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perception models that distinguish walkable sidewalks from adjacent unsafe regions.

By Hakan Calim, Anamaria Dumitrescu, Adarsh Bhandary Panambur, Huzaifa Asif, Andreas Maier
arXiv Machine Learning
Jul 24

Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility

arXiv:2607. 21444v1 Announce Type: cross Abstract: Reliable electric vehicle (EV) charging infrastructure is a cornerstone of sustainable, low-carbon cities, yet urban climate stress such as extreme heat, heavy precipitation, and humidity increasingly raises equipment fault risk and undermines the resilience of urban energy and mobility services.

By Cande Lian (School of Management, Foshan University, Foshan, China), Wentao Zeng (School of Management, Foshan University, Foshan, China), Jiabin Wu (School of Management, Foshan University, Foshan, China), Yiming Bie (School of Transportation, Jilin University, Changchun, China), Wei Zhou (Department of Civil and Environmental Engineering, National University of Singapore)
arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv Machine Learning
Jul 24

Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

arXiv:2607. 21068v1 Announce Type: new Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep learning (DL) models have been widely utilized.

By Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
arXiv AI
Jul 24

Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis

arXiv:2607. 20691v1 Announce Type: cross Abstract: Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical imaging, their trustworthiness is often limited by the quality and granularity of available supervision.

By Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah, Afzel Noore