arXiv:2608. 14472v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise.
By Abhishek Shukla, Ankur Sinha, Faiz Hamid
arXiv:2606. 29582v1 Announce Type: cross Abstract: Bilevel optimization has become an influential and widely adopted framework for addressing hierarchical optimization problems in machine learning, providing an effective approach to modeling the interaction between two levels of optimization, with applications such as hyperparameter tuning, meta-learning, adversarial training, and data poisoning.
By Abhishek Shukla, Ankur Sinha, Faiz Hamid
arXiv:2603. 15106v2 Announce Type: replace Abstract: Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately.
By Mark Deutel, Simon Geis, Axel Plinge
ONNX-Net introduces a universal representation for neural architectures using natural language descriptions, enabling instant performance prediction across diverse search spaces. The authors present ONNX-Bench, a benchmark of over 600k architecture–accuracy pairs compiled from open‑source NAS‑bench networks in ONNX format. Experiments demonstrate strong zero‑shot predictive performance with minimal pretraining, overcoming the limitations of cell‑based, graph‑encoded approaches.
By Shiwen Qin, Alexander Auras, Shay B. Cohen, Elliot J. Crowley, Michael Moeller, Linus Ericsson, Jovita Lukasik
arXiv:2607. 15745v1 Announce Type: new Abstract: Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches.
By Anxhelo Shehu, Enes Stastoli, Arben Cela
The paper compares genetic algorithm (GA) and gradient descent (GD) training for a distance‑encoding biomorphic‑informational neural network (DEBI‑NN) designed for low‑data medical datasets. A spatial backpropagation scheme was implemented for GD, and both optimizers were evaluated on synthetic, radiomic, and fetal cardiotocography datasets. Across all experiments, GA consistently outperformed GD, achieving higher classification accuracy and more stable decision boundaries, while GD struggled with the interdependent spatial parameters of DEBI‑NN.
By Amine Boukhari, Boglarka Ecsedi, Laszlo Papp, Mathieu Hatt