arXiv Machine Learning
Jun 5

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang
arXiv AI
Sep 7

Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

Large language models (LLMs) are increasingly used in high‑stakes real‑world systems such as financial markets. This study demonstrates that enhancing individual LLM capability can actually worsen system‑level outcomes by making models behave more similarly, leading to correlated actions that increase risk. Using an agent‑based simulation of LLM traders, the authors show that while higher capability can reduce market risk when reasoning is accurate, it can amplify risk when agents share misinformation, revealing a capability paradox.

By Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo