arXiv Machine Learning

Matched Queries for Curvature and Density at Branching Junctions

arXiv:2609. 01319v1 Announce Type: cross Abstract: At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center.

Hugging Face Trending Papers
Jul 5

Asymptotic-Preserving A Posteriori Analysis of Diffusion and Flow-Matching Samplers

Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.

arXiv Machine Learning
Sep 10

Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training

The paper investigates how hard‑ReLU training behaves when perturbations have a finite radius. It shows that the usual infinitesimal sensitivities are insufficient to predict the response at a chosen radius, and it characterizes the intermediate regime where the perturbation radius scales with the gradient‑descent step. The authors derive crossing indices, a uniform endpoint expansion for separated transverse events, and provide explicit remainder terms in contractive affine regions to certify finite candidate comparisons, supported by experiments on nonlinear networks.

By Xiaoyang Li, Runni Zhou, Xinghao Yan
arXiv Machine Learning
Sep 11

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale

EGGROLL replaces dense Gaussian perturbations in evolution strategies with low‑rank Gaussian products, enabling practical optimization of large language models while maintaining exactness on quadratic objectives. The paper analyzes the mean update field, error bounds, and shows that rank‑one perturbations add only a small variance penalty compared to dense ES. A new leave‑one‑out estimator, LOO‑ROLL, further reduces estimator MSE and improves post‑training performance on transformer blocks and GSM8K accuracy.

By Ege C. Kaya, Abolfazl Hashemi