Hugging Face Trending Papers

Predicting Deep Neural Network Training Outcomes from Early Training Telemetry

Read the original on Hugging Face Trending Papers →

Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 11

ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks

ExpTest is an autonomous learning‑rate controller that uses the training loss curve as an online signal to perform sequential statistical tests on theoretically motivated windows, detecting convergent behavior and triggering learning‑rate reductions. It combines a covariance‑based initial learning‑rate estimate, curvature‑motivated window sizing, and a two‑phase test‑driven decay, relying on the approximately exponential decay predicted under linearized network dynamics. Experiments on regression, classification, forecasting, and natural‑language tasks across various architectures show that ExpTest achieves competitive performance compared to hand‑tuned SGD baselines and recent learning‑rate‑free methods, without requiring manual initial learning‑rate selection or predefined scheduling.

By Zan Chaudhry, Naoko Mizuno
arXiv AI
Aug 10

Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction

arXiv:2608. 06993v1 Announce Type: cross Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack.

By Gregor Molan (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Grafika Jati (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Francesco Barchi (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Andrea Acquaviva (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Alja\v{z} Osterman (LE-Tehnika d.o.o., \v{S}uceva 27, Kranj, 4000, Slovenia), Martin Molan (Comtrade AI GmbH, Grafenauweg 8, Zug, 6300, Switzerland)