← Back to all news
arXiv Machine Learning September 22, 2026 By Jihwan Kim, Dogyoon Song, Chulhee Yun

Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 18

Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression

arXiv:2605. 24316v3 Announce Type: replace Abstract: Mini-batching is central to large-scale optimization, yet its role in statistical scaling laws remains limited.

By Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou
safety
More like this →
arXiv Machine Learning
Jul 15

Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps

arXiv:2607. 12360v1 Announce Type: new Abstract: The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others.

By Subham Singh, Ashutosh Mishra, Subha Raut
More like this →
arXiv Machine Learning
Sep 15

Optimal Learning Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay

arXiv:2602.06797v3 Announce Type: replace-cross Abstract: We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training d...

By Binghui Li, Zilin Wang, Fengling Chen, Shiyang Zhao, Ruiheng Zheng, Lei Wu
More like this →
arXiv Machine Learning
Jun 3

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

arXiv:2603. 05691v3 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models.

By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
safety
More like this →
arXiv AI
Jul 28

Scale Weight Decay and Train Better

arXiv:2607. 23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data.

By Anuj Apte
safety
More like this →
arXiv AI
Jun 3

Sign Lock-In: Randomly Initialized Weight Signs Persist and Bottleneck Sub-Bit Model Compression

arXiv:2602. 17063v2 Announce Type: replace-cross Abstract: Sub-bit model compression targets storage below one bit per weight; as magnitudes are aggressively compressed, the sign bit becomes a fixed-cost bottleneck.

By Akira Sakai, Yuma Ichikawa
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea