arXiv:2604. 17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions.
By Zikun Ye, Hema Yoganarasimhan
arXiv:2608.22582v1 Announce Type: cross
Abstract: Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including de...
By Julia Romberg, Tobias Gummer, Gabriella Lapesa, Tanja Kunz, Claudia Wagner
arXiv:2606. 12332v1 Announce Type: cross Abstract: Evaluating multi-turn dialogue is challenging because quality emerges across turns rather than within individual responses.
By Paul He, Shiva Kasiviswanathan, Dominik Janzing
The paper introduces ProSE, a framework for AI assistants that generate proposals while considering users’ bounded rationality and evaluability constraints. It proposes a KL‑regularised bounded‑rational binary response model and a depth‑2 Bayes‑adaptive planner, “ProSE‑Plan,” which scores proposals by expected responses and resulting belief updates. Experiments on graph simulations show that “ProSE‑Plan” outperforms evaluability‑unaware and myopic baselines, especially when evaluation cost is high, and that informative probes are crucial for effective assistance.
By Yifan Zhu, Sammie Katt, Samuel Kaski
arXiv:2601. 07055v2 Announce Type: replace Abstract: As high-quality data becomes increasingly difficult to obtain, self-evolution without curated training data has emerged as a promising paradigm.
By Zhenrui Yue, Kartikeya Upasani, Xianjun Yang, Suyu Ge, Shaoliang Nie, Yuning Mao, Zhe Liu, Dong Wang
arXiv:2608.29856v1 Announce Type: new
Abstract: Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during...
By Yifan Chen, Haitao Li, Qingyao Ai, Fengbin Zhu, Tat-Seng Chua, Min Zhang, Yiqun Liu