arXiv Machine Learning

LESS: Lightweight Evolutionary Supernet Search in Minutes

LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.

arXiv AI
Sep 17

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

The paper presents a Pareto atlas of LLM inference optimizations, mapping cost, quality, and latency trade‑offs for Qwen2.5‑7B‑Instruct on L4, A100, and H100 GPUs. Using 54 measured configurations and a calibrated simulator, it identifies 18 of 36 setups on the Pareto frontier, showing that combined methods outperform single ones. Quality tests reveal that AWQ 4bit and FP8 weights offer significant latency reductions while largely preserving accuracy, but naive FP8 KV caching fails to answer any questions correctly.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv AI
6d ago

ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers

ENAS is a hardware‑aware neural architecture search framework tailored for TinyML on microcontrollers. It uses a static feasibility check, a cell‑based search space with various block types and skip connections, and a three‑stage hybrid search strategy (random → top‑K → mutation) with cross‑run caching. The framework runs efficiently without GPUs, achieving significant search‑time speedups and competitive accuracy on Visual Wake Words and Melanoma Cancer benchmarks across a range of microcontrollers.

By Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan
arXiv Machine Learning
4d ago

Scaling Zero-Order Pretraining through Model Sharding

arXiv:2609.37899v1 Announce Type: new Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...

By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv Machine Learning
Sep 22

Are Coreset Selection Methods Worth Their Cost?

The paper evaluates coreset selection methods by incorporating both selection and training time into a unified wall‑clock budget, using a standardized benchmark across four datasets and multiple selectors. Across numerous budget anchors, simple random or full‑data training consistently outperforms sophisticated selectors, and selection costs are dominated by a full‑dataset scan that cannot be amortized. The study also identifies when subset reuse can justify selection and reports several correctness fixes in a popular codebase.

By Yangze Liu, Zhongyi Han
arXiv AI
6d ago

Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study

The paper introduces TGL-NSGA-II, a low‑fidelity framework that uses a pretrained teacher to stratify samples by difficulty and class, then applies a short knowledge‑distillation step (KD‑Lite) before scoring candidates on a stratified evaluation set. The teacher‑guided scores are fused with a Gaussian‑process surrogate to select candidates for full evaluation, and the method is evaluated on keyword spotting and bird‑call classification tasks. Results show high Kendall‑τ values (0.74 and 0.62), a 41% reduction in proxy‑score variance, and improved hypervolume and false‑positive rates compared to full NSGA‑II, while running 2.2× faster under a constrained evaluation budget.

By Soumen Garai, Suman Samui
arXiv Machine Learning
Jul 14

HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NAS

arXiv:2607. 11746v1 Announce Type: new Abstract: With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential.

By Shambhavi Balamuthu Sampath, Behzad Shomali, Nael Fasfous, Moritz Thoma, Judeson Anthony Fernando, Lukas Frickenstein, Pierpaolo Mori, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele
arXiv Machine Learning
Aug 28

Curating Same-Family Neural Networks for LLM-Guided Model Improvement: A Controlled Case Study

The study investigates whether a curated same-family neural network experiment can guide large language model (LLM)-based improvements for a low-performing target model under equal generation and evaluation budgets. Using TuneNNGen, an extension of NNGPT, the authors compare source-guided generation with target-only generation on CIFAR-10, SVHN, Imagenette, and CIFAR-100 datasets, achieving significant accuracy gains across these benchmarks. The results demonstrate that the benefits depend on source-target compatibility and LLM adaptation, rather than merely on stored source accuracy.

By Kabir Dev Paul Baghel, Radu Timofte, Dmitry Ignatov