The study tests deliberately biased AI assistants and finds that such bias improves human performance on tasks like misinformation evaluation, financial investment, and graduate education compared to neutral AI. However, participants undervalue biased AI and overvalue neutral AI, even when performance is similar. When two AI biases flank a participant’s perspective, performance gains are maintained while reducing the perceived cost and one‑sided influence.
By Shiyang Lai, Jiwoong Choi, Junsol Kim, Nadav Kunievsky, Yujin Potter, James Evans
arXiv:2608. 11794v1 Announce Type: cross Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement.
By Adrian Rauchfleisch, Andreas Jungherr
The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.
By Nirav Patel, Emily Wenger, Christopher Buccafusco
The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain.
The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.
By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv:2604. 08525v2 Announce Type: replace Abstract: Large language models (LLMs) are trained to align with user preferences through methods like reinforcement learning.
By Addison J. Wu, Ryan Liu, Shuyue Stella Li, Yulia Tsvetkov, Thomas L. Griffiths
arXiv:2604.22230v2 Announce Type: replace-cross
Abstract: Performance manipulation arises when agents exploit easily measurable, routine tasks to inflate observable outcomes without contributing genu...
By Xiaoyun Qiu, Yang Yu, Haifeng Xu
arXiv:2608.29803v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
By Lin Chen, Yitong Chen, Yong Li
arXiv:2509.08494v2 Announce Type: replace-cross
Abstract: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures....
By Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis
The study reports a $2 imes2$ randomized field experiment involving 1,072 calls to a German Santa Claus telephone hotline, with 89 child conversations (median age 6) meeting inclusion criteria. Children were routed to one of four LLM voice agents that varied in persona (Santa vs. Helper) and framing (persuasive nudges toward prosocial wishes vs. neutral). Persuasive framing increased the likelihood of a prosocial wish from 11.6% to 45.7%, while persona authority had little effect on compliance but did influence engagement, with children more likely to hang up on the Helper within the first minute.
By Thilo Tamme (Technical University of Munich), David Steck (Technical University of Munich), Anton Hantel (Massachusetts Institute of Technology)
arXiv:2608.24748v1 Announce Type: cross
Abstract: How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dynamics of AI through an open-ended, mixe...
By Jacy Reese Anthis, Erik Brynjolfsson, James Evans