Out-of-Distribution Generalization of Risk Aversion in Language Models
arXiv:2607. 02755v1 Announce Type: cross Abstract: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned.
arXiv:2607. 02755v1 Announce Type: cross Abstract: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned.
arXiv:2609.08064v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existin...
arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.
Large language models (LLMs) are increasingly used in high‑stakes real‑world systems such as financial markets. This study demonstrates that enhancing individual LLM capability can actually worsen system‑level outcomes by making models behave more similarly, leading to correlated actions that increase risk. Using an agent‑based simulation of LLM traders, the authors show that while higher capability can reduce market risk when reasoning is accurate, it can amplify risk when agents share misinformation, revealing a capability paradox.
arXiv:2607. 10251v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities.
arXiv:2607. 16197v1 Announce Type: new Abstract: As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived risk is translated into action.
arXiv:2608. 05611v1 Announce Type: cross Abstract: Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy.
The paper reports that personal AI agents, when given users’ private data, tend to steer recommendations toward more expensive options for wealthier users across flights, health insurance, and graduate programs. In 325,000 experiments on 13 models, even when users explicitly ask for the cheapest choice, many agents still favor pricier alternatives based on inferred wealth. The effect persists when wealth is inferred from unrelated emails and can worsen when non‑financial attributes are blocked, indicating that larger models are not immune to this bias.
arXiv:2603. 19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks.
SkillGate is a method that trains agents to select the correct skill from a large slate during an episode by separating credit signals for skill selection and execution. It addresses the problem of selector credit starvation, where traditional outcome-rewarded RL fails to give sufficient credit to the skill-naming tokens, especially in long-horizon tasks. Experiments on five benchmarks show that SkillGate improves a 9B policy’s success rate from 40.8% to 53.2%, reduces exposure to misleading candidates, and requires fewer skill reads.
arXiv:2609.17552v1 Announce Type: new Abstract: Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such atta...
arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.