AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,618 stories · RSS feed

arXiv AI
Aug 11

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

arXiv:2608. 01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR).

By Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
arXiv Machine Learning
Aug 11

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

arXiv:2608. 07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv Machine Learning
Aug 11

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv:2608. 07935v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student.

By Meilin Yang (Renmin University of China, Beijing, China), Zixuan Ding (Renmin University of China, Beijing, China), Jianhao Nie (Renmin University of China, Beijing, China), Weite Zhang (Renmin University of China, Beijing, China), Yuxin Zhang (Renmin University of China, Beijing, China), Zhiming Shao (Renmin University of China, Beijing, China), Li Yu (Renmin University of China, Beijing, China), Zhe Fu (Renmin University of China, Beijing, China)
arXiv Machine Learning
Aug 11

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

arXiv:2608. 08008v1 Announce Type: new Abstract: Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning.

By Ibne Farabi Shihab, Fariya Afrin
arXiv Machine Learning
Aug 11

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

By Andrea Combette, Nelly Pustelnik, Antoine Venaille
arXiv Machine Learning
Aug 11

Can Graph Learning Learn Circuits?

arXiv:2608. 08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient to reproduce a particular behavior.

By Chester Tan, Moritz Lampert, Courtney Maynard, Ankit Ramakrishnan, Tina Eliassi-Rad, Ingo Scholtes
arXiv Machine Learning
Aug 11

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

arXiv:2608. 08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE.

By Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou
arXiv Machine Learning
Aug 11

F2STNet: Fair and Federated Spectral-Temporal Modeling for Graph Forecasting

arXiv:2608. 09082v1 Announce Type: new Abstract: Spatiotemporal prediction on graph-structured data is central to traffic forecasting and environmental monitoring, yet decentralized and heterogeneous data complicate both sequence modeling and collaborative training.

By Jiayi Zhang, Jinfeng Xu, Hewei Wang, Siyuan Cen, Haidong Huang, Yiyao Zhan, Zheyu Chen, Jinjiang You, Ai Jian, Edith C. H. Ngai
arXiv Machine Learning
Aug 11

Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research

arXiv:2608. 07522v1 Announce Type: cross Abstract: We present a structured review of commonly used Explainable machine learning (XML) methodologies, including global and local interpretability tools such as SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) plots.

By Krishna Padmanabhan, Minxin Lu, Dai Feng, Natalia KanDobrosky, Sai Konduri, Heather J. Litman, Achilleas Livieratos
arXiv Machine Learning
Aug 11

Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting

arXiv:2608. 07643v1 Announce Type: cross Abstract: Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing in practice: a detector already trained on the class one wants to count.

By Lucas Gouveia Omena Lopes, William W. M. Lira, Alexandre M. Lima, Thales M. A. Vieira