arXiv:2602. 16745v2 Announce Type: replace-cross Abstract: Test-time scaling can improve model performance by aggregating stochastic reasoning trajectories.
By Zhangyi Liu, Huaizhi Qu, Xiaowei Yin, He Sun, Yanjun Han, Tianlong Chen, Zhun Deng
arXiv:2608.30805v1 Announce Type: cross
Abstract: Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregat...
By Jos\'e Mar\'ia Lago, Albert Castellana, Edgars Nem\v{s}e
arXiv:2606. 08360v1 Announce Type: cross Abstract: Peer-referral recruitment systems such as respondent-driven sampling are critical for studying and intervening on hidden populations affected by infectious diseases.
By Lingkai Kong, Hezi Jiang, Andrew Ma, Keyu Wang, Akseli Kangaslahti, Milind Tambe
The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.
By Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen
arXiv:2605.25200v3 Announce Type: replace
Abstract: Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single us...
By Xiang Cheng, Yulan Hu, Lulu Zheng, Xiangwen Zhang, Zheng Pan, Xin Li, Yong Liu
arXiv:2608. 07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator.
By Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang
CityReal is a modular framework that uses large language model agents to simulate human-aligned urban behavior. It models agents as intention-driven decision makers who pursue coherent mobility and activity plans, learning habits and preferences over time. By training textual adapters to align agent decisions with observed population statistics, CityReal improves realism at both micro and macro levels and can scale to tens of thousands of agents for analyzing crowd density, place popularity, mobility flows, and well‑being under various urban scenarios.
By Nicolas Bougie, Xiaotong Ye, Narimasa Watanabe
The paper investigates whether complex communication topologies are necessary for effective multi‑agent debate (MAD) among large language models. It demonstrates that a simple random-without-replacement routing policy—where each agent debates with two newly sampled peers each round—consistently improves the accuracy‑cost trade‑off in sparse MAD setups. Additionally, the study shows that lightweight deliberation stopping can further reduce inference costs without sacrificing accuracy.
By Boxuan Wang, Zhuoyun Li, Xiaowei Huang, Yi Dong
The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.
By Taehyung Kim, Jongeun Choi
arXiv:2608. 20316v1 Announce Type: new Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost.
By Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
arXiv:2607. 11632v1 Announce Type: new Abstract: Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality.
By Jiangtao Han, Shoufeng Ma, Shuxian Xu, Geng Li, Shuai Ling, Ning Jia, Zhengbing He
The paper presents a method to align a language‑model‑based crowd agent with aggregate mobility data by fine‑tuning it to match observed destination compositions derived from origin‑to‑destination flows. The approach uses iterative proportional fitting to reweight the model’s destination distribution and corrects for dominant destination inflation by training a low‑rank adapter on resampled trajectories. Experiments on mobile network counts from two baseball games show a 25% reduction in destination‑share error while maintaining similar grid correlation across policies.