arXiv AI

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

arXiv:2607. 10251v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities.

arXiv AI
Jun 10

Superficial Beliefs in LLM Decision-Making

arXiv:2606. 11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure.

By Gabriel Freedman, Francesca Toni
arXiv AI
Sep 25

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.

By Igor Bogdanov, Olga Manakina, Chung-Horng Lung
arXiv AI
Jul 28

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

By Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
arXiv AI
2d ago

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

The paper introduces TRACE, a token‑level objective designed to reduce multi‑turn safety risks in large language models. TRACE assigns each token a weight based on the discounted return of a refusal‑attributable advantage, comparing a frozen reference model with a refusal‑ablated copy to credit early tokens for later refusal evidence. Evaluated across five open‑weight models and seven multi‑turn attacks, TRACE achieves the lowest attack success rate in all 35 model‑attack pairs while maintaining model utility within 1.23 points on MMLU and HellaSwag.

By Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang