arXiv Machine Learning

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using anchor samples and an ExtraTrees residual model. The method achieves high Spearman correlations across multiple NAS benchmarks and selects architectures that approach the best‑known accuracy with only about 1% of the training cost, without relying on fully trained labels for ranking. It demonstrates robust cross‑space performance and improves over static capacity proxies, especially in size‑only search spaces.

Hugging Face Trending Papers
Sep 10

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using an ExtraTrees model. The method achieves high Spearman correlations across multiple benchmark spaces and selects architectures with accuracy close to the ground‑truth best, all while using only about 1% of the training cost of fully training the candidate set. CoRA‑NAS provides a single configuration that works across different search spaces, combining cross‑space ranking robustness with efficient architecture selection.

arXiv Machine Learning
Sep 14

RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search

RiPPLE is a method for ranking neural architectures across an entire search space using only a small fraction of early training data. It treats partial training as labels for a limited set of anchor architectures, extrapolates their learning curves, and propagates these surrogate labels to other architectures without requiring per‑candidate features. The approach is evaluated on twelve benchmark cells from four search‑space families and the larger DARTS space, demonstrating its effectiveness in ranking quality, label efficiency, and architecture selection.

By Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang
arXiv Machine Learning
1d ago

Prediction-powered Neural Architecture Search

The paper introduces PPNAS, a prediction‑powered inference method for neural architecture search that combines a small set of accurately evaluated architectures with a large set of zero‑cost proxy (ZCP) evaluations. PPNAS uses the ordinal information from ZCPs to generate pairwise ranking supervision and applies a debiasing step to reconcile discrepancies between proxy and true performance rankings. Experiments show that PPNAS achieves state‑of‑the‑art results in predictor‑based NAS under limited evaluation budgets.

By Pascal Janetzky, Yuxin Wang, Michael Klar, Stefan Feuerriegel
arXiv Machine Learning
Sep 1

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

PRIME is a plug‑in residual input‑conditioned mixture of experts that preserves the original dense prediction path while adding low‑rank, input‑dependent logit corrections. By initializing residuals to zero, PRIME matches the baseline dense model at training start and stabilizes conditional estimation with multi‑bag aggregation and EMA load biases. Experiments on Avazu and Criteo across 13 CTR architectures show modest AUC and LogLoss gains, with PRIME outperforming APG on FiBiNET and DCNv2 while using fewer parameters and lower latency.

By Heng Yao, Siyun Hou, Tianying Liu, Yulou Shu, Yong He, Chuan Yuan, Kaibin Qiu, Guowei Chen, Jiayu Zhao, Chao Yu, Ke Ding
arXiv Machine Learning
Sep 22

Efficient K-generalizable Learned Search

arXiv:2603.06159v2 Announce Type: replace-cross Abstract: Learned top-K search improves the accuracy-latency trade-off of graph-based vector search, but existing methods are designed for a fixed K: s...

By Yifan Peng, Jiafei Fan, Xingda Wei, Sijie Shen, Rong Chen, Jianning Wang, Xiaojian Luo, Wenyuan Yu, Jingren Zhou, Haibo Chen
arXiv Machine Learning
1d ago

LESS: Lightweight Evolutionary Supernet Search in Minutes

LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.

By Aviral Gandhi, Jinglue Xu, Jialong Li, Hitoshi Iba
arXiv AI
Jul 21

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.

By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu