arXiv Machine Learning By Tiexin Ding

Weibull Weight-Scale Parameter Evolution under AdamW Training Dynamics

Read the original on arXiv Machine Learning →

arXiv:2606. 19367v1 Announce Type: new Abstract: Building on a two-parameter Weibull framework for diagnosing transformer weight distributions, we study why the Weibull weight-scale parameter $\lambda$ grows, overshoots, and then relaxes during AdamW training.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.