arXiv:2607. 24662v1 Announce Type: new Abstract: Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap.
By Tianpeng Li, Xuan Guo, Wenjun Wang, Wang Zhang, Pengfei Jiao
arXiv:2607. 03436v1 Announce Type: new Abstract: Routing among large language models (LLMs) promises better quality at lower cost, motivated by the reported gap between learned routers and a per-instance oracle.
By Teng-Ruei Chen
arXiv:2607. 04113v1 Announce Type: new Abstract: Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $\sigma_{\min}$, at which the score is stiff and the flow develops a boundary layer.
By Shiheng Zhang
Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.
By Barada Sahu (Cabal AI), Shivesh Pandey (Para AI)
arXiv:2604. 18194v2 Announce Type: replace Abstract: Single-step generators promise high-fidelity synthesis at a fraction of the inference and training cost of ordinary differential equation (ODE)-based flow models, a central concern when compute is limited.
By Arkadii Kazanskii, Tatiana Petrova, Andrey Ustyuzhanin, Konstantin Bagrianskii, Aleksandr Puzikov, Radu State
arXiv:2608. 00675v1 Announce Type: cross Abstract: Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against.
By Alexander Scheinker
arXiv:2606. 32005v1 Announce Type: cross Abstract: Stochastic Gradient Descent ($\textsf{SGD}$) is one of the most classical optimization algorithms with favorable theoretical guarantees, yet the practical implementation of $\textsf{SGD}$ differs subtly from its well-known form and is often referred to as Shuffling Stochastic Gradient Descent ($\textsf{Shuffling SGD}$).
By Zijian Liu
arXiv:2605. 15219v3 Announce Type: replace Abstract: Can AI systems discover new knowledge through iterative self-improvement, and at what cost?
By Salman Avestimehr, Ken Duffy, Muriel M\'edard
arXiv:2607. 16811v4 Announce Type: replace Abstract: Drift detectors that work tend not to explain themselves, and drift detectors that explain themselves tend to fail in high dimension.
By Behnam Asadi
arXiv:2606. 16541v1 Announce Type: new Abstract: Autoformalization, translating natural-language mathematics into formal proof assistants, is bottlenecked not by translation fluency but by \emph{faithfulness}: a formal statement can typecheck and be provable, yet still encode a different theorem than the source intended.
By Noor Islam S. Mohammad, Tamim Sheikh
arXiv:2606. 14488v1 Announce Type: cross Abstract: Recent finite-time analyses of nonlinear two-time-scale stochastic approximation show that under contractive assumptions the slow iterate $Y_k$ with stepsizes $\beta_k=\Theta(k^{-1})$ and $\alpha_k=\Theta(k^{-a})$, $a\in(1/2,1)$, generally satisfies a mean-square rate of order $k^{-a}$; decoupled $k^{-1}$ rates require strong local linearity.
By Dhruv Sarkar, Vaneet Aggarwal