arXiv AI By Zhanliang Zhu, Ziwei Li, Yuchen Liu, Liujun Zhu, Ruiqi Wu, Tongqing Shen, Junliang Jin, Jianyun Zhang

The average-farmer illusion in language-model simulations of agricultural decisions

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 25

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

The article introduces the Artificial Societies Benchmark, a validation framework designed to evaluate synthetic populations used in research. It comprises eleven tests covering internal, construct, and external validity, drawing on twenty human data sources and comparing nine language models. The benchmark links specific research uses to the evidence required and assesses how results vary with different respondent information, revealing that strong performance in one domain does not guarantee fidelity in others.

By Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He