arXiv Machine Learning

RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search

RiPPLE is a method for ranking neural architectures across an entire search space using only a small fraction of early training data. It treats partial training as labels for a limited set of anchor architectures, extrapolates their learning curves, and propagates these surrogate labels to other architectures without requiring per‑candidate features. The approach is evaluated on twelve benchmark cells from four search‑space families and the larger DARTS space, demonstrating its effectiveness in ranking quality, label efficiency, and architecture selection.

Hugging Face Trending Papers
Sep 10

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using an ExtraTrees model. The method achieves high Spearman correlations across multiple benchmark spaces and selects architectures with accuracy close to the ground‑truth best, all while using only about 1% of the training cost of fully training the candidate set. CoRA‑NAS provides a single configuration that works across different search spaces, combining cross‑space ranking robustness with efficient architecture selection.

arXiv Machine Learning
Sep 11

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using anchor samples and an ExtraTrees residual model. The method achieves high Spearman correlations across multiple NAS benchmarks and selects architectures that approach the best‑known accuracy with only about 1% of the training cost, without relying on fully trained labels for ranking. It demonstrates robust cross‑space performance and improves over static capacity proxies, especially in size‑only search spaces.

By Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang
arXiv Machine Learning
1d ago

LESS: Lightweight Evolutionary Supernet Search in Minutes

LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.

By Aviral Gandhi, Jinglue Xu, Jialong Li, Hitoshi Iba
arXiv Machine Learning
Aug 27

ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures

ONNX-Net introduces a universal representation for neural architectures using natural language descriptions, enabling instant performance prediction across diverse search spaces. The authors present ONNX-Bench, a benchmark of over 600k architecture–accuracy pairs compiled from open‑source NAS‑bench networks in ONNX format. Experiments demonstrate strong zero‑shot predictive performance with minimal pretraining, overcoming the limitations of cell‑based, graph‑encoded approaches.

By Shiwen Qin, Alexander Auras, Shay B. Cohen, Elliot J. Crowley, Michael Moeller, Linus Ericsson, Jovita Lukasik
arXiv Machine Learning
Sep 1

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

PRIME is a plug‑in residual input‑conditioned mixture of experts that preserves the original dense prediction path while adding low‑rank, input‑dependent logit corrections. By initializing residuals to zero, PRIME matches the baseline dense model at training start and stabilizes conditional estimation with multi‑bag aggregation and EMA load biases. Experiments on Avazu and Criteo across 13 CTR architectures show modest AUC and LogLoss gains, with PRIME outperforming APG on FiBiNET and DCNv2 while using fewer parameters and lower latency.

By Heng Yao, Siyun Hou, Tianying Liu, Yulou Shu, Yong He, Chuan Yuan, Kaibin Qiu, Guowei Chen, Jiayu Zhao, Chao Yu, Ke Ding
arXiv AI
Aug 20

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.

By Tanay Chowdhury, Saeideh Shahrokh Esfahani
arXiv Computation and Language
3d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv Machine Learning
Aug 27

Why and When Neural Networks Improve Local Approximation in Optimization

The paper investigates why neural network surrogates sometimes improve and sometimes worsen derivative‑free optimisation performance. It identifies three key factors—role (whether the surrogate proposes candidates or replaces gradients), radius (the neighbourhood within which a local model is reliable), and room (whether the base method can still progress)—that determine when a learned local model is beneficial. Experiments on 117 benchmark instances show that providing surrogate‑approved candidates boosts success rates, while replacing gradients or ignoring the radius can reduce them.

By Chengkuo Bian, Pengcheng Xie