arXiv AI

Designing Compact Neural Architectures via Neuron Gating and Mixed Activation

arXiv:2608. 14443v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss.

arXiv AI
Jun 30

Bilevel Optimization for Neural Architecture Search

arXiv:2606. 29582v1 Announce Type: cross Abstract: Bilevel optimization has become an influential and widely adopted framework for addressing hierarchical optimization problems in machine learning, providing an effective approach to modeling the interaction between two levels of optimization, with applications such as hyperparameter tuning, meta-learning, adversarial training, and data poisoning.

By Abhishek Shukla, Ankur Sinha, Faiz Hamid
arXiv Machine Learning
Aug 27

ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures

ONNX-Net introduces a universal representation for neural architectures using natural language descriptions, enabling instant performance prediction across diverse search spaces. The authors present ONNX-Bench, a benchmark of over 600k architecture–accuracy pairs compiled from open‑source NAS‑bench networks in ONNX format. Experiments demonstrate strong zero‑shot predictive performance with minimal pretraining, overcoming the limitations of cell‑based, graph‑encoded approaches.

By Shiwen Qin, Alexander Auras, Shay B. Cohen, Elliot J. Crowley, Michael Moeller, Linus Ericsson, Jovita Lukasik
arXiv Machine Learning
Sep 18

Genetic algorithm vs. gradient descent for training a neural network architecture dedicated to low data regimes in small medical datasets

The paper compares genetic algorithm (GA) and gradient descent (GD) training for a distance‑encoding biomorphic‑informational neural network (DEBI‑NN) designed for low‑data medical datasets. A spatial backpropagation scheme was implemented for GD, and both optimizers were evaluated on synthetic, radiomic, and fetal cardiotocography datasets. Across all experiments, GA consistently outperformed GD, achieving higher classification accuracy and more stable decision boundaries, while GD struggled with the interdependent spatial parameters of DEBI‑NN.

By Amine Boukhari, Boglarka Ecsedi, Laszlo Papp, Mathieu Hatt
arXiv Machine Learning
Sep 24

NGN: Learning Neural Network Size as a Differentiable Count

The paper introduces the Neurogenesis Network (NGN), a differentiable framework that learns the optimal number of ordered structural components in a neural network during training. By using a learnable boundary to select an active prefix of components, NGN can grow from a compact initialization and later discard unused parts. Experiments across MLPs, CNNs, GNNs, Transformers, state‑space models, LoRA, and adapters show that the learned prefixes perform comparably to fixed‑size models, demonstrating that structural capacity can be optimized directly as a count.

By Lixing Li