AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,632 stories · RSS feed

arXiv Machine Learning
Jun 24

Lightweight Test-Time Adaptation for EMG-Based Gesture Recognition

arXiv:2601. 04181v2 Announce Type: replace Abstract: Reliable long-term decoding of gestures from surface electromyography (EMG) is hindered by signal drift caused by electrode displacement, muscle fatigue, and/or posture changes.

By Nia Touko, Matthew O A Ellis, Cristiano Capone, Alessio Burrello, Elisa Donati, Luca Manneschi
arXiv AI
Jun 24

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

arXiv:2606. 23927v1 Announce Type: new Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities.

By Yarin Yerushalmi Levi, Roy Betser, Amit Giloni, Lidor Erez, Itay Gershon, Oren Rachmil, Sindhu Padakandla, Roman Vainshtein
arXiv Machine Learning
Jun 24

Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making

arXiv:2508. 16420v3 Announce Type: replace Abstract: Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especially when the target return lies in underrepresented regions of the dataset.

By Yue Pei, Hongming Zhang, Chao Gao, Martin M\"uller, Yingying Zhang, Mengxiao Zhu, Hao Sheng, Ziliang Chen, Liang Lin, Haogang Zhu
arXiv AI
Jun 24

A Survey on Federated Causal Discovery and Inference

arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.

By Xianjie Guo, Yuwei Wang, Guodu Xiang, Xiaoli Tang, Kui Yu, Han Yu, Qiang Yang