arXiv:2602. 16745v2 Announce Type: replace-cross Abstract: Test-time scaling can improve model performance by aggregating stochastic reasoning trajectories.
By Zhangyi Liu, Huaizhi Qu, Xiaowei Yin, He Sun, Yanjun Han, Tianlong Chen, Zhun Deng
arXiv:2608.30805v1 Announce Type: cross
Abstract: Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregat...
By Jos\'e Mar\'ia Lago, Albert Castellana, Edgars Nem\v{s}e
arXiv:2606. 08360v1 Announce Type: cross Abstract: Peer-referral recruitment systems such as respondent-driven sampling are critical for studying and intervening on hidden populations affected by infectious diseases.
By Lingkai Kong, Hezi Jiang, Andrew Ma, Keyu Wang, Akseli Kangaslahti, Milind Tambe
The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.
By Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen
arXiv:2605.25200v3 Announce Type: replace
Abstract: Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single us...
By Xiang Cheng, Yulan Hu, Lulu Zheng, Xiangwen Zhang, Zheng Pan, Xin Li, Yong Liu
arXiv:2608. 07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator.
By Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang