Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents
arXiv:2607. 03015v1 Announce Type: new Abstract: Forecasting future events has attracted growing attention as a testbed for general-purpose AI.
arXiv:2607. 06166v1 Announce Type: new Abstract: Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events.
arXiv:2607. 03015v1 Announce Type: new Abstract: Forecasting future events has attracted growing attention as a testbed for general-purpose AI.
arXiv:2607. 00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecasters, since a proper scoring rule such as the Brier score is computed from outcomes alone and is minimized in expectation by the true probability.
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...
arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.
arXiv:2607. 16229v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.
arXiv:2606. 04342v1 Announce Type: cross Abstract: Multi-step time series forecasting (MSF) is commonly evaluated using point-wise error metrics such as mean squared error (MSE), implicitly treating the conditional mean as a sufficient target.
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
The paper investigates the coherence of probabilistic forecasts produced by language models, particularly in the context of life‑decision support. Using a de Finetti‑based method, the authors elicit forecasts for events derived from stock return data and compute the maximum Dutch‑book profit via linear programming, which quantifies incoherence. They find significant incoherence, especially when events have complex logical relationships or when irrelevant context is present, and suggest that alternative training strategies could improve coherence.
arXiv:2609.36061v1 Announce Type: new Abstract: In quantitative finance, standard regression losses are misaligned with the economics of return prediction. As the conditional mean of financial log-re...
The paper introduces loss‑conditioned state execution, a model‑agnostic technique that decides whether to apply a world model’s proposed state change or keep the current state based on whether the change reduces downstream loss. It formalizes state movability as the existence of a loss‑reducing feasible correction and constructs loss‑specific proposals from predictive distributions, executing them only when a groupwise lower confidence bound on loss improvement is positive. Experiments on forecasting and dynamics benchmarks show that the method accepts updates for a subset of cases, achieving lower bounded loss than persistence or always executing the proposal, and highlights that event predictability and loss‑based decisions must be evaluated separately.
UQ-LOB is a lightweight, encoder‑agnostic module that adds uncertainty quantification to any pretrained limit order book (LOB) encoder. It offers two variants: UQ‑regression, which outputs a calibrated Gaussian over future tick displacement, and UQ‑classification, which outputs a categorical distribution over down/up/stationary. On 5.2 billion LOB events across seven cryptocurrency assets, UQ‑regression achieves near‑nominal 68 % interval coverage, and selecting the top 10 % most confident predictions boosts directional macro F1 by 0.11–0.15 for regression and 0.05–0.11 for classification, reaching F1 scores of 0.88 (down) and 0.83 (up) at a 5‑second horizon.