arXiv AI By Pei Wang, Xu Chen, Ji-Rong Wen

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

Read the original on arXiv AI →

arXiv:2608. 19689v1 Announce Type: new Abstract: LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 2

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani