AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,527 stories · RSS feed

arXiv Machine Learning
Jul 16

Power Homotopy for Zeroth-Order Non-Convex Optimizations

arXiv:2511. 13592v2 Announce Type: replace-cross Abstract: The existing method of GS-PowerOpt solves the non-convex optimization problem of the form $\max_{\boldsymbol{x} \in \mathbb{R}^d} f(\boldsymbol{x})$ through maximizing a Gaussian-smoothed surrogate $F_{N,\sigma}(\boldsymbol{\mu}) = \mathbb{E}_{\boldsymbol{x}\sim\mathcal{N}(\boldsymbol{\mu},\sigma^2 I_d)}[e^{N f(\boldsymbol{x})}]$.

By Chen Xu
arXiv AI
Jul 16

AIMO Interpretability Challenge

arXiv:2607. 13899v1 Announce Type: new Abstract: We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms.

By Michal \v{S}tef\'anik, Philipp Mondorf, Andreas Waldis, Qianying Liu, Chuan Yang, Michal Spiegel, Josef Kucha\v{r}, Marek Kadl\v{c}\'ik, Adam Vawda-Oomerjee, Chaoran Liu, Simon Frieder, Barbara Plank, Fazl Barez, Pontus Stenetorp
arXiv AI
Jul 16

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

arXiv:2601. 02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora.

By Amirali Ebrahimzadeh, Seyyed M. Salili
arXiv AI
Jul 16

Music-to-Dance Generation via Atomic Movements

arXiv:2607. 13978v1 Announce Type: cross Abstract: Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music.

By Xinhao Cai, Yixuan Sun, Minghang Zheng, Qingchao Chen, Xin Jin, Song-chun Zhu, Yang Liu
arXiv Machine Learning
Jul 16

Value Drifts: Tracing Value Alignment During LLM Post-Training

arXiv:2510. 26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems.

By Mehar Bhatia, Shravan Nayak, Gaurav Kamath, Marius Mosbach, Karolina Sta\'nczak, Vered Shwartz, Siva Reddy