Hugging Face Trending Papers

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using an ExtraTrees model. The method achieves high Spearman correlations across multiple benchmark spaces and selects architectures with accuracy close to the ground‑truth best, all while using only about 1% of the training cost of fully training the candidate set. CoRA‑NAS provides a single configuration that works across different search spaces, combining cross‑space ranking robustness with efficient architecture selection.

arXiv Machine Learning
Sep 11

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS is a two‑stage neural architecture search framework that first uses a static coarse ranking (CoRA‑Rank) based on capacity and structure‑at‑initialization proxies, then refines this ranking with low‑cost learning‑curve extrapolation (CoRA‑Refine) using anchor samples and an ExtraTrees residual model. The method achieves high Spearman correlations across multiple NAS benchmarks and selects architectures that approach the best‑known accuracy with only about 1% of the training cost, without relying on fully trained labels for ranking. It demonstrates robust cross‑space performance and improves over static capacity proxies, especially in size‑only search spaces.

By Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang
arXiv Machine Learning
Sep 14

RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search

RiPPLE is a method for ranking neural architectures across an entire search space using only a small fraction of early training data. It treats partial training as labels for a limited set of anchor architectures, extrapolates their learning curves, and propagates these surrogate labels to other architectures without requiring per‑candidate features. The approach is evaluated on twelve benchmark cells from four search‑space families and the larger DARTS space, demonstrating its effectiveness in ranking quality, label efficiency, and architecture selection.

By Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang
arXiv Machine Learning
1d ago

Prediction-powered Neural Architecture Search

The paper introduces PPNAS, a prediction‑powered inference method for neural architecture search that combines a small set of accurately evaluated architectures with a large set of zero‑cost proxy (ZCP) evaluations. PPNAS uses the ordinal information from ZCPs to generate pairwise ranking supervision and applies a debiasing step to reconcile discrepancies between proxy and true performance rankings. Experiments show that PPNAS achieves state‑of‑the‑art results in predictor‑based NAS under limited evaluation budgets.

By Pascal Janetzky, Yuxin Wang, Michael Klar, Stefan Feuerriegel
arXiv AI
Jul 21

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.

By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu
arXiv Machine Learning
Sep 22

Efficient K-generalizable Learned Search

arXiv:2603.06159v2 Announce Type: replace-cross Abstract: Learned top-K search improves the accuracy-latency trade-off of graph-based vector search, but existing methods are designed for a fixed K: s...

By Yifan Peng, Jiafei Fan, Xingda Wei, Sijie Shen, Rong Chen, Jianning Wang, Xiaojian Luo, Wenyuan Yu, Jingren Zhou, Haibo Chen
arXiv Machine Learning
Sep 22

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B. whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."

By Chenyu Zhu, Ruoyu Zhao, Zhichao Lu
arXiv Machine Learning
1d ago

LESS: Lightweight Evolutionary Supernet Search in Minutes

LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.

By Aviral Gandhi, Jinglue Xu, Jialong Li, Hitoshi Iba