arXiv Machine Learning

CAGE-NAS: Certified Functional Descent for Efficient Model Growth

CAGE-NAS introduces a method for deciding when to grow a neural network by evaluating an admissibility criterion on approximations of the functional gradient. If the current architecture allows a certified Functional Gradient Descent step, it remains unchanged; otherwise, a function-preserving expansion is performed and re-evaluated. Using a tangent-space-based instance with regularized projection, the approach achieves architectures that rank above the 99.8th percentile in held‑out RMSE among all admissible alternatives within the same parameter budget, without enumerating them during growth.

arXiv Machine Learning
Jun 16

Functional Gradient Descent with Adaptive Representations

arXiv:2606. 16926v1 Announce Type: cross Abstract: Functional optimization problems are typically solved by optimizing the parameters of a fixed representation, such as a neural network, resulting in highly nonconvex losses that complicate both training and theoretical analysis.

By Daniel Csillag, Rodrigo Schuller, Pedro Dall'Antonia, Leonidas Guibas, Luiz Velho, Tiago Novello
arXiv Machine Learning
Sep 24

A lift for input-convex neural net training

The paper introduces the "lift" technique for training input‑convex neural networks, replacing the traditional non‑negative weight constraint enforced by projected gradient descent or a softplus map. By adding a learnable slack variable and an unconstrained network that processes a permutation‑invariant batch summary, the lift couples batch‑dependent latent weights to the gradient, increasing update variance and enabling faster escape from the softplus shoulder. Experiments show that when the softplus method stalls at the shoulder, the lift achieves tighter fits and reconstructs targets roughly three times faster, while both methods agree when the shoulder is rarely reached.

By Ali Siahkoohi
arXiv AI
Jul 21

Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs

arXiv:2607. 16568v1 Announce Type: new Abstract: Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program.

By Abdallah Khemais (ISITCOM, University of Sousse)