arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
arXiv:2608.30373v1 Announce Type: new
Abstract: Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. Howev...
By Minsoo Song, Chanwoo Kim, Sugyeong Eo, Chanjun Park
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
By Umesh Bodhwani, Thanh Tran, Kai Wei
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
arXiv:2608.10503v2 Announce Type: replace
Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. T...
By Davood Wadi, Mohsen Ghodrat, Matthew Philp
Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. However, for subjective rubric-based scoring, inter-ag...
arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
By Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, Max Pellert
arXiv:2606. 18258v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries.
By Sunnie S. Y. Kim, Margit Bowler, Leon A Gatys
arXiv:2608. 05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings.
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
arXiv:2606. 12730v1 Announce Type: new Abstract: Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior.
By Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez
arXiv:2608.29803v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
By Lin Chen, Yitong Chen, Yong Li