arXiv Machine Learning

Adapt Only When It Pays: Budgeted Decision-Loss Priority for Delayed Online Time-Series Adaptation

arXiv:2606. 25068v1 Announce Type: new Abstract: Online time-series forecasters receive labels only after horizon-dependent delays, while every adaptation step spends limited compute.

arXiv Machine Learning
Aug 11

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.

By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
arXiv Machine Learning
Sep 2

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

The paper investigates when online adaptation benefits edge time‑series forecasting under distribution drift, using a leakage‑free streaming protocol on six public multivariate datasets. It shows that the warmup budget for static baselines and the choice of learning rate can bias perceived adaptation gains, and that a validation‑only procedure selecting warmup and optimizer rates yields Adam outperforming SGD with momentum in most settings. The study also examines accuracy versus adaptation‑state memory and per‑update latency for different adaptation strategies, highlighting parameter‑efficient variants that are nondominated on the memory axis.

By Takumi Fujimoto, Hiroaki Nishi
arXiv Machine Learning
Jul 30

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.

By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv Machine Learning
Aug 19

Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts

Dynamic Regime-Aware Conformal Prediction (DRACP) is a new method that blends density‑ratio estimation, localized kernel weighting, and probabilistic regime‑aware weighting with a self‑tuning online significance controller to produce reliable prediction intervals under multiple distribution shifts. The authors prove finite‑sample validity with oracle weights, provide a coverage‑gap bound for estimated weights, and give deterministic or regret guarantees for the online controller. In experiments on 48 real forecasting series—including euro‑area inflation, US macroeconomic and energy indicators, and daily financial data—DRACP achieves the most reliable calibration, maintaining coverage close to the nominal 0.90 and never falling below 0.80, while other methods achieve narrower intervals but with higher under‑coverage. whyItMatters":"DRACP offers a principled trade‑off between calibration and efficiency, ensuring that prediction intervals meet coverage standards even when economic data exhibit covariate shift, concept drift, and latent regimes."

By Bogdan Oancea