AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,515 stories · RSS feed

arXiv AI
Jul 17

SeeSE3: Emergence of 3D Space in Vision Features

arXiv:2607. 14228v1 Announce Type: cross Abstract: In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space.

By Caroline Chen, Sayna Ebrahimi, Fedor Kitashov, Ming-Hsuan Yang, Leonidas Guibas, Viorica P\u{a}tr\u{a}ucean, Maks Ovsjanikov
arXiv AI
Jul 17

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

arXiv:2607. 14353v1 Announce Type: cross Abstract: As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems.

By Joshua A. Kroll, Andrew Smart, R. Stuart Geiger, Abigail Z. Jacobs
arXiv AI
Jul 17

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

arXiv:2607. 14816v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias.

By Saima Afrin, Alessandro Midolo, Camilo Escobar-Vel\'asquez, Mario Linares-V\'asquez, Weiyuan Ding, Bowen Xu, Massimiliano Di Penta, Antonio Mastropaolo
arXiv Machine Learning
Jul 17

What Do Temporal Graph Learning Models Learn?

arXiv:2510. 09416v4 Announce Type: replace Abstract: Learning on temporal graphs has become a central topic in graph representation learning, with numerous benchmarks indicating the strong performance of state-of-the-art models.

By Abigail J. Hayes, Tobias Schumacher, Markus Strohmaier
arXiv AI
Jul 17

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv:2406. 01514v4 Announce Type: replace-cross Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback.

By Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu