The study investigates sycophancy in Chinese large language models (LLMs) by analyzing 364,941 responses from DeepSeek, Qwen, and Doubao to 12,165 yes/no factual questions derived from real-world search queries. It examines how user beliefs, reasoning, and anti-sycophancy prompts affect the distribution of correct, incorrect, and uncertain answers, finding that anti-sycophancy instructions can reduce belief-aligned errors but often increase uncertainty. The results show that preventing agreement with false beliefs does not necessarily preserve factual accuracy, underscoring the need for transition-level evaluation in Chinese-language factual QA.
By Geng Liu, Feng Li, Mengxiao Zhu, Francesco Pierri
arXiv:2607. 11039v1 Announce Type: cross Abstract: Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities.
By Pengping Tan, Baoquan Zhao, Zhenhui Peng
The paper introduces JobMate, a persona‑grounded conversational agent that transforms peer job‑seeking posts into interactive dialogues to aid career exploration. In a study with 24 participants, JobMate helped users select relevant cases, ask follow‑up questions, and articulate constraints, contrasting with a static browsing tool that left reconstruction to users. The authors discuss design implications for blending authentic peer experiences with generative AI.
By Pengping Tan, Baoquan Zhao, Shuai Ma, Zhenhui Peng
arXiv:2607. 25620v1 Announce Type: new Abstract: Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted.
By Federico Cabitza, Gianluca Colombo
arXiv:2608.29803v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
By Lin Chen, Yitong Chen, Yong Li
arXiv:2606. 05890v1 Announce Type: cross Abstract: LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display?
By Salvatore Greco, Hainiu Xu, Jacopo Domenicucci, Yulan He, Sylvie Delacroix
arXiv:2609.16436v1 Announce Type: cross
Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the...
By Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj
arXiv:2604.19139v4 Announce Type: replace-cross
Abstract: Repeated praise, canned reassurance, familiar contrasts, and conspicuous vocabulary are recurring subjects in discussions of large language m...
By Shuai Wu, Xue Li, Zhijun Wang, Bolun Liu, Weilin Cai, Zihao Su, Ran Wang
The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv:2609.14638v1 Announce Type: cross
Abstract: This paper is an encore submission of our 2026 journal article "Expertise and Information Seeking in the Age of Generative AI: New Procedures, New Pr...
By Alexi Orchard, Shannon Lodoen
arXiv:2606. 22748v2 Announce Type: replace-cross Abstract: Some professional authors are beginning to use AI tools to help produce their fiction writing.
By Neel Gupta, Maria Antoniak, Melanie Walsh
arXiv:2603. 00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants?
By Jason Miklian, Kristian Hoelscher, John E. Katsos