The paper introduces a score‑calibrated robustness framework that transforms any fixed point predictor into a decision‑relevant uncertainty representation using distribution‑free conformal calibration. By employing the conformal score as the core unit of robustness, the authors derive both reliability‑based robust optimization and target‑oriented Conformal Robust Satisficing formulations, linking them through a shared robust decision frontier and a fragility measure. Experiments on synthetic data and a real online‑grocery inventory case study demonstrate the framework’s ability to improve reliability, reduce costs, and provide interpretable uncertainty scales for black‑box predictors.
By Lingjie Zhao, Hansheng Jiang, Wei Qi
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
Dynamic Regime-Aware Conformal Prediction (DRACP) is a new method that blends density‑ratio estimation, localized kernel weighting, and probabilistic regime‑aware weighting with a self‑tuning online significance controller to produce reliable prediction intervals under multiple distribution shifts. The authors prove finite‑sample validity with oracle weights, provide a coverage‑gap bound for estimated weights, and give deterministic or regret guarantees for the online controller. In experiments on 48 real forecasting series—including euro‑area inflation, US macroeconomic and energy indicators, and daily financial data—DRACP achieves the most reliable calibration, maintaining coverage close to the nominal 0.90 and never falling below 0.80, while other methods achieve narrower intervals but with higher under‑coverage.
whyItMatters":"DRACP offers a principled trade‑off between calibration and efficiency, ensuring that prediction intervals meet coverage standards even when economic data exhibit covariate shift, concept drift, and latent regimes."
By Bogdan Oancea
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree.
arXiv:2606. 25068v1 Announce Type: new Abstract: Online time-series forecasters receive labels only after horizon-dependent delays, while every adaptation step spends limited compute.
By Xibai Wang
arXiv:2607. 16229v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.
By Rishab Ghosh, Vinay Devarakonda
arXiv:2609.13840v1 Announce Type: new
Abstract: A contract-logistics spare-parts operator is paid on order-level service: an order counts only if every requested line is fulfilled, yet forecasters ar...
By Joo Ern Chin, Shih-Fen Cheng, Aldy Gunawan
The paper introduces loss‑conditioned state execution, a model‑agnostic technique that decides whether to apply a world model’s proposed state change or keep the current state based on whether the change reduces downstream loss. It formalizes state movability as the existence of a loss‑reducing feasible correction and constructs loss‑specific proposals from predictive distributions, executing them only when a groupwise lower confidence bound on loss improvement is positive. Experiments on forecasting and dynamics benchmarks show that the method accepts updates for a subset of cases, achieving lower bounded loss than persistence or always executing the proposal, and highlights that event predictability and loss‑based decisions must be evaluated separately.
By Jintao Xu, Zhengyu Chen, Ben Zhang, Yongzhi Qi, Jianshen Zhang
arXiv:2609.24559v1 Announce Type: new
Abstract: We present $t_0$, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: $\texttt{t0-a...
By Lucas Meyer, Claudio Sole, Huikan Xiang, Nicolas Li, Lucas Franceschino, Arnau Quera-Bofarull, Maarten P. Scholl, Joachim Fainberg, Geoffrey N\'egiar
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
arXiv:2607. 13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed.
By Sagar Deb, Ashwanth Krishnan
arXiv:2607. 11653v1 Announce Type: new Abstract: Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management.
By Ivane Antonov, Sohom Mukherjee, Richard Pibernik, Yo Joong Choe