AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,195 stories · RSS feed

arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv Machine Learning
Jul 24

From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python

arXiv:2607. 21069v1 Announce Type: new Abstract: The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal.

By Muntasir Adnan, Manile Srun, Carlos C. N. Kuhn
arXiv AI
Jul 24

From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics

arXiv:2607. 21327v1 Announce Type: cross Abstract: Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems.

By Muhsen Hammoud
arXiv Machine Learning
Jul 24

SPECTRA: State-Space Exogenous Context and Temporal-Frequency Resolution Architecture for Probabilistic Energy Forecasting

arXiv:2607. 20587v1 Announce Type: cross Abstract: Modern power systems increasingly require probabilistic forecasts amid interacting uncertainties from renewable intermittency, flexible demand, market volatility, and weather-dependent generation.

By Hang Ye, Xinyan Jiang, Yuedong Shi, Yangxin Zhu, Jianming Wei, Tian Zheng, Xiaoying Zheng, Yongxin Zhu
arXiv AI
Jul 24

Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving

arXiv:2508. 16947v2 Announce Type: replace-cross Abstract: Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency biased behaviors, overlooking the inherent behavioral diversity of human driving.

By Fan Ding, Xuewen Luo, Fucai Ke, Hwa Hui Tew, Susilawati Susilawati, Vishnu Monn Baskaran, Junn Yong Loo
arXiv AI
Jul 24

Hybrid MKNF with Classical Negation in the Rule Component

arXiv:2607. 21202v1 Announce Type: cross Abstract: Hybrid MKNF knowledge bases under the well-founded semantics integrate Description Logics with Logic Programming.

By Arun Raveendran Nair Sheela (Universit\'e Clermont Auvergne, LIMOS Laboratory, Thales), Christophe Rey (Universit\'e Clermont Auvergne, LIMOS, CNRS, France), Florence De Grancey (Thales)
arXiv AI
Jul 24

Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis

arXiv:2607. 20691v1 Announce Type: cross Abstract: Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical imaging, their trustworthiness is often limited by the quality and granularity of available supervision.

By Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah, Afzel Noore
arXiv Machine Learning
Jul 24

How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals

arXiv:2501. 02211v3 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied.

By Messi H. J. Lee