arXiv Machine Learning

L-SR1: Learned Symmetric-Rank-One Preconditioning

arXiv:2508. 12270v3 Announce Type: replace Abstract: End-to-end deep learning has achieved impressive results but often relies on large labeled datasets, exhibits limited generalization to unseen scenarios, and incurs substantial computational cost.

arXiv Machine Learning
Jun 19

Deep-Unfolded Coordination

arXiv:2606. 19920v1 Announce Type: cross Abstract: Distributed optimization is a highly scalable and structurally transparent technique to solve multi-agent robotics problems; however, such methods often suffer from the need for highly-specialized, problem-specific hyperparameter tunings.

By Hunter Kuperman, Minchan Jung, Rahul V. Ghosh, Alex Oshin, Evangelos A. Theodorou
arXiv Machine Learning
5d ago

Gradient Surgery for Physics-Informed Neural Networks

The paper introduces PAM-GS, a physics-aware gradient surgery technique for Physics-Informed Neural Networks (PINNs). It addresses the highly imbalanced multi-task optimisation problem in PINNs by adaptively mitigating task interference based on observed gradient conflicts. Experiments on four PDE benchmarks show that PAM-GS achieves competitive solution accuracy while maintaining strong task-balanced performance, outperforming existing methods on most problems.

By Thomas Borsani, Giuseppe Di Fatta
arXiv Computer Vision
Sep 24

Learn2Splat: Extending the Horizon of Learned 3DGS Optimization

arXiv:2605.15760v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) optimization is most commonly performed using general-purpose first-order optimizers such as Adam or SGD. Although rob...

By Naama Pearl, Stefano Esposito, Haofei Xu, Amit Peleg, Patricia Gschossmann, Lorenzo Porzi, Peter Kontschieder, Gerard Pons-Moll, Andreas Geiger
arXiv Machine Learning
1d ago

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.

By Cristian McGee, El Houcine Bergou, Aritra Dutta