AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,629 stories · RSS feed

arXiv Machine Learning
Aug 10

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

arXiv:2608. 06429v1 Announce Type: cross Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior.

By Yong Yang, Roger Newman-Norlund, Xiang Guan, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Sophie Arheix-Parras, Srihari Nelakuditi, Leonardo Bonilha, Christopher Rorden, Rutvik H. Desai, Julius Fridriksson
arXiv Machine Learning
Aug 10

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

arXiv:2608. 06674v1 Announce Type: cross Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems.

By Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando, Harshala Gammulle, Basura Fernando, Sanka Rasnayake, A V Subramanyam, Sridha Sridharan, Clinton Fookes