At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recove...
arXiv:2608.22334v1 Announce Type: new
Abstract: Near a smooth data manifold, one tangent space summarizes local geometry. At a branch point, the corresponding first-order object is instead a measure...
By Ziqi Zhao, Qingjian Ni
arXiv:2607. 04113v1 Announce Type: new Abstract: Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $\sigma_{\min}$, at which the score is stiff and the flow develops a boundary layer.
By Shiheng Zhang
Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.
arXiv:2608. 13922v1 Announce Type: new Abstract: Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified.
By Guoqing Zhang, Zhaixin Chen
The paper investigates how hard‑ReLU training behaves when perturbations have a finite radius. It shows that the usual infinitesimal sensitivities are insufficient to predict the response at a chosen radius, and it characterizes the intermediate regime where the perturbation radius scales with the gradient‑descent step. The authors derive crossing indices, a uniform endpoint expansion for separated transverse events, and provide explicit remainder terms in contractive affine regions to certify finite candidate comparisons, supported by experiments on nonlinear networks.
By Xiaoyang Li, Runni Zhou, Xinghao Yan
arXiv:2607. 05872v1 Announce Type: new Abstract: Memory-efficient optimizers such as GaLore train large language models by projecting gradients onto a rank-r subspace recomputed every T steps, assuming this subspace is a slowly drifting object that can be tracked.
By Noel Thomas
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
arXiv:2605. 23434v2 Announce Type: replace Abstract: Approximate inference over inducing variables is the central computational bottleneck of Deep Gaussian Processes (DGPs).
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2608. 01547v1 Announce Type: cross Abstract: Drifting objectives compare a target and model distribution through a vector field observed noisily at finitely many locations.
By Sam Andersson, Ricky Mol\'en
arXiv:2607. 04113v2 Announce Type: replace Abstract: Diffusion and Gaussian-interpolant flow-matching samplers approach data through a terminal noise floor $\varepsilon$, a singular limit for manifold-supported or rank-deficient data.
By Shiheng Zhang
EGGROLL replaces dense Gaussian perturbations in evolution strategies with low‑rank Gaussian products, enabling practical optimization of large language models while maintaining exactness on quadratic objectives. The paper analyzes the mean update field, error bounds, and shows that rank‑one perturbations add only a small variance penalty compared to dense ES. A new leave‑one‑out estimator, LOO‑ROLL, further reduces estimator MSE and improves post‑training performance on transformer blocks and GSM8K accuracy.
By Ege C. Kaya, Abolfazl Hashemi