Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations.
The paper proves that in time‑series validation three desirable properties—training sufficiency, test coverage, and temporal causality—cannot all be satisfied simultaneously. It introduces quantitative bounds involving the smallest training fraction (α), test coverage (β), future training fraction (Λ), and distance to nearest future training point (δ), showing that exceeding the causal frontier α+β=1 requires training on future data that must lie within (1−α)T of a test point. The authors demonstrate that the impact of such future leakage depends on distance rather than amount, and compare different validation schemes (walk‑forward, k‑fold, purged k‑fold) in terms of their position on this Pareto frontier, illustrating the trade‑offs with empirical results on noise data.
By Jiayu Li
arXiv:2607. 04113v1 Announce Type: new Abstract: Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $\sigma_{\min}$, at which the score is stiff and the flow develops a boundary layer.
By Shiheng Zhang
arXiv:2607. 03436v1 Announce Type: new Abstract: Routing among large language models (LLMs) promises better quality at lower cost, motivated by the reported gap between learned routers and a per-instance oracle.
By Teng-Ruei Chen
arXiv:2609. 10954v1 Announce Type: new Abstract: Continual world models must decide whether new data justify changing the model.
By Anqi Peter Li, Kaden Kim
Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.
arXiv:2608. 00675v1 Announce Type: cross Abstract: Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against.
By Alexander Scheinker
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.
By Barada Sahu (Cabal AI), Shivesh Pandey (Para AI)
arXiv:2609.23378v1 Announce Type: new
Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first...
By Zijiang Yang
arXiv:2604. 18194v2 Announce Type: replace Abstract: Single-step generators promise high-fidelity synthesis at a fraction of the inference and training cost of ordinary differential equation (ODE)-based flow models, a central concern when compute is limited.
By Arkadii Kazanskii, Tatiana Petrova, Andrey Ustyuzhanin, Konstantin Bagrianskii, Aleksandr Puzikov, Radu State
The paper derives precise cost formulas for self‑calibrating monitors that adjust thresholds online to maintain a specified long‑run false‑alarm rate under arbitrary drift. It shows that the guarantee is an accounting identity, independent of the monitored signal, and provides exact evidence identities for both step and ramp drift scenarios, as well as an exact law for the fluctuation of the certificate’s own alarm rate. Additionally, it proves that any monitor designed to tolerate a drift class is blind to all faults in the difference of that class, identifying the blind set for speed‑bounded drift classes and quantifying power outside this set with a sharp Gaussian projection bound.
By Abdou-Raouf Atarmla
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
By Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi