The paper introduces a score‑calibrated robustness framework that transforms any fixed point predictor into a decision‑relevant uncertainty representation using distribution‑free conformal calibration. By employing the conformal score as the core unit of robustness, the authors derive both reliability‑based robust optimization and target‑oriented Conformal Robust Satisficing formulations, linking them through a shared robust decision frontier and a fragility measure. Experiments on synthetic data and a real online‑grocery inventory case study demonstrate the framework’s ability to improve reliability, reduce costs, and provide interpretable uncertainty scales for black‑box predictors.
By Lingjie Zhao, Hansheng Jiang, Wei Qi
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
Dynamic Regime-Aware Conformal Prediction (DRACP) is a new method that blends density‑ratio estimation, localized kernel weighting, and probabilistic regime‑aware weighting with a self‑tuning online significance controller to produce reliable prediction intervals under multiple distribution shifts. The authors prove finite‑sample validity with oracle weights, provide a coverage‑gap bound for estimated weights, and give deterministic or regret guarantees for the online controller. In experiments on 48 real forecasting series—including euro‑area inflation, US macroeconomic and energy indicators, and daily financial data—DRACP achieves the most reliable calibration, maintaining coverage close to the nominal 0.90 and never falling below 0.80, while other methods achieve narrower intervals but with higher under‑coverage.
whyItMatters":"DRACP offers a principled trade‑off between calibration and efficiency, ensuring that prediction intervals meet coverage standards even when economic data exhibit covariate shift, concept drift, and latent regimes."
By Bogdan Oancea
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree.
arXiv:2606. 25068v1 Announce Type: new Abstract: Online time-series forecasters receive labels only after horizon-dependent delays, while every adaptation step spends limited compute.
By Xibai Wang
arXiv:2607. 16229v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.
By Rishab Ghosh, Vinay Devarakonda