Hugging Face Trending Papers

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry

Multi-agent systems are increasingly used for forecasting future events, as deliberation among multiple LLMs is believed to improve reasoning and calibration. Yet existing approaches overlook a critical design choice: what information each agent receives.

arXiv AI
Sep 2

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting) is a new approach that reorganizes how evidence is used in LLM-based forecasting systems. Instead of a monolithic prediction that aggregates all evidence at once, LEAP examines each evidence item separately, elicits likelihood parameters, and combines them with an explicit prior to produce a posterior distribution. The method supports continuous, single-choice, and multi-choice forecasts and has been shown to improve prediction and calibration metrics across models on a benchmark covering forecasting, information-seeking, and browsing tasks.

By Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
arXiv AI
Sep 17

Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making

The paper investigates whether in-context learning (ICL) in large language model agents reflects genuine recursive reasoning or simply statistical extrapolation. By testing LLM agents in a public goods game with manipulated historical feedback, the authors compare decision quality to a rational expectations equilibrium benchmark. They find that disrupting historical patterns eliminates the benefits of longer context, especially in highly interdependent settings, indicating that ICL behavior aligns more with statistical extrapolation than strategic reasoning.

By Yu Liu, Wenwen Li, Yifan Dou, Guangnan Ye
arXiv AI
Sep 25

When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.

By Yufeng Wang