arXiv:2607. 03998v1 Announce Type: new Abstract: The local sharpness of the loss, the top Hessian eigenvalue $\lambda_1$, determines the largest stable gradient step, but measuring it normally requires Lanczos or Hessian-vector iterations.
By Ashmitha R, J\"org Frochte
The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.
By Rub\'en Dar\'io Guerrero
arXiv:2606. 29119v1 Announce Type: cross Abstract: We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle outer loop over neural-network parameters or structure is worth building.
By Ramchand Kumaresan
arXiv:2605. 01928v2 Announce Type: replace Abstract: We optimize losses that jump: spiking thresholds, quantized layers, and discrete routing put jumps in the forward pass, where backpropagation does not apply.
By An T. Le
arXiv:2608. 06825v1 Announce Type: new Abstract: Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward.
By Pahan Dewasurendra
arXiv:2608. 08103v1 Announce Type: new Abstract: Smooth acyclicity constraints answer whether a weighted support is a DAG, whereas structure learning asks which support change should be made.
By Rui Wu, Zongyuan Chen, Hong Xie
The paper proposes constraining the query and key projection matrices in Transformer attention to the Stiefel manifold and optimizing them with a Riemannian Adam optimizer. It demonstrates that this geometric constraint yields significant performance gains on a CIFAR‑10 patch benchmark, with the constrained model outperforming standard AdamW by up to +6.79 percentage points. The authors also show that weight decay has no effect on the constrained frames and that the improvement is driven by a scale‑free step size rather than the manifold projection or equivariance properties.
By Rub\'en Dar\'io Guerrero
arXiv:2609.26272v1 Announce Type: new
Abstract: Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to...
By Jian Xu
arXiv:2606. 14640v1 Announce Type: new Abstract: We study Online Convex Optimization (OCO) over a convex set $K\subseteq \mathbb R^d$, where in each round $t$ the learner selects $x_t\in K$ and then observes a convex loss $f_t:K\to[0,1]$, with the goal of minimizing regret to the best fixed decision in hindsight.
By Simone Di Gregorio, Anupam Gupta, Stefano Leonardi, Matteo Russo
The paper shows that Worst‑Case Optimal Recovery (OR) and Bayesian learning solve the same Gaussian‑quadratic‑Hilbert problems, linking the radius of information to a nugget‑optimized Gaussian process posterior variance. It evaluates three Bayesian systems, demonstrating that OR can outperform Bayesian methods in certain calibration and reproducibility metrics, yet split‑conformal and other approaches can beat OR in interval scoring, especially under covariate shift. The authors propose matching the guarantee tool to the data regime and auditing that regime first.
By Gordei Verbii
arXiv:2608. 08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose coordinate and inspection time are chosen only after the data are seen.
By Ibne Farabi Shihab, Adria Binte Habib
arXiv:2608.23094v1 Announce Type: new
Abstract: One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity...
By Gordei Verbii