arXiv:2607. 16197v1 Announce Type: new Abstract: As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived risk is translated into action.
By Bowen Sun, Rui Min, Yuxi Wang, Brian Odegaard, Qi Wang, Jing Du
arXiv:2609.08064v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existin...
By Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan zheng
arXiv:2608. 01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma.
By Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The Anh Han
arXiv:2607. 02755v1 Announce Type: cross Abstract: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned.
By Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie, Andrew Lin, Abhitej Bokka, Elliott Thornley
arXiv:2606. 11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure.
By Gabriel Freedman, Francesca Toni
The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.
By Igor Bogdanov, Olga Manakina, Chung-Horng Lung
arXiv:2606. 07552v1 Announce Type: cross Abstract: Large language models exhibit innate behavioral tendencies when deployed as strategic agents -- notably a risk-averse "turtle" bias toward defensive play.
By Augustin Chan
arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.
By Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
arXiv:2607. 24765v1 Announce Type: cross Abstract: Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context.
By Gi-Hun Lee, Joong Yull Park
arXiv:2608.18265v2 Announce Type: replace-cross
Abstract: We introduce a general, easy-to-implement AI-based method for modeling and analyzing the structure and complexity of human behavior. We assig...
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
The paper introduces TRACE, a token‑level objective designed to reduce multi‑turn safety risks in large language models. TRACE assigns each token a weight based on the discounted return of a refusal‑attributable advantage, comparing a frozen reference model with a refusal‑ablated copy to credit early tokens for later refusal evidence. Evaluated across five open‑weight models and seven multi‑turn attacks, TRACE achieves the lowest attack success rate in all 35 model‑attack pairs while maintaining model utility within 1.23 points on MMLU and HellaSwag.
By Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang