arXiv AI By Zeyu Tang, Sang T. Truong, Deonna Owens, Shreyas Sharma, Yibo Jacky Zhang, Brando Miranda, Sanmi Koyejo

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

Read the original on arXiv AI →

arXiv:2605. 12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv Computation and Language
Sep 23

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.

By Alexandre Cristov\~ao Maiorano