arXiv:2607. 21999v1 Announce Type: new Abstract: Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape.
By Jiaxin Deng, Junbiao Pang
arXiv:2607. 18343v1 Announce Type: cross Abstract: Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor.
By Radhakrishna Achanta, Will Reed
Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. In particular, the perturbation radius $ρ$ is typically treated as an isolated tuning parameter, despite defining the neighborhood in which SAM measures sharpness.
arXiv:2607. 20235v1 Announce Type: cross Abstract: Classical regularization removes the binary-collision singularity from the Kepler problem, but its value as a representation for learned Hamiltonian dynamics has not been systematically isolated.
By Abhishek Shankar
arXiv:2609.30342v1 Announce Type: new
Abstract: iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as A...
By Rajit Rajpal, Benedict Leimkuhler
arXiv:2606. 04279v1 Announce Type: new Abstract: Machine-learned (ML) exchange-correlation (XC) functionals aim to replace human-designed density functional approximations by learning directly from reference data, but they still do not consistently outperform traditional $\mathcal{O}(N^4)$-scaling hybrid functionals.
By Eike S. Eberhard, Luca A. Thiede, Abdul Aldossary, Andreas Burger, Nicholas Gao, Vignesh Bhethanabotla, Al\'an Aspuru-Guzik, Stephan G\"unnemann
arXiv:2608. 03197v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear.
By Jiaxin Deng, Junbiao Pang
arXiv:2609. 30271v1 Announce Type: new Abstract: Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate.
By Gongyue Zhang, Honghai Liu
arXiv:2606. 13657v2 Announce Type: replace Abstract: On-policy distillation (\textsc{OPD}) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student trajectories and dense teacher supervision.
By Guo Yu, Wenlin Liu, Yulan Hu, Hao-Xuan Ma, Jun-Peng Jiang, Han-Jia Ye
The paper introduces the "lift" technique for training input‑convex neural networks, replacing the traditional non‑negative weight constraint enforced by projected gradient descent or a softplus map. By adding a learnable slack variable and an unconstrained network that processes a permutation‑invariant batch summary, the lift couples batch‑dependent latent weights to the gradient, increasing update variance and enabling faster escape from the softplus shoulder. Experiments show that when the softplus method stalls at the shoulder, the lift achieves tighter fits and reconstructs targets roughly three times faster, while both methods agree when the shoulder is rarely reached.
By Ali Siahkoohi
arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.
By Ravi Satya Durga Prasad Yenugula
arXiv:2606. 11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training.
By Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren