arXiv AI

The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network

arXiv:2508. 21380v3 Announce Type: replace-cross Abstract: Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents.

arXiv Machine Learning
Jun 25

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

arXiv:2606. 26094v1 Announce Type: new Abstract: For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention.

By Babak Rahmani, Sebastian Dziadzio, Joschka Str\"uber, Sergio Hern\'andez-Guti\'errez, Matthias Bethge
arXiv Machine Learning
Jul 7

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

arXiv:2506. 07468v4 Announce Type: replace Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities.

By Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques