← Back to all news
arXiv Machine Learning August 26, 2026 By Tiexin Ding

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 19

Weibull Weight-Scale Parameter Evolution under AdamW Training Dynamics

arXiv:2606. 19367v1 Announce Type: new Abstract: Building on a two-parameter Weibull framework for diagnosing transformer weight distributions, we study why the Weibull weight-scale parameter $\lambda$ grows, overshoots, and then relaxes during AdamW training.

By Tiexin Ding
llmssafety
More like this →
arXiv Machine Learning
Jul 27

Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

arXiv:2607. 21866v1 Announce Type: new Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells.

By Kaihua Ding
benchmarks
More like this →
arXiv AI
Jun 11

Unifying Learning Dynamics and Generalization in Transformers Scaling Law

arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.

By Chiwun Yang
llms
More like this →
arXiv AI
Jul 28

Scale Weight Decay and Train Better

arXiv:2607. 23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data.

By Anuj Apte
safety
More like this →
arXiv AI
Jul 15

Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG

arXiv:2607. 11950v1 Announce Type: cross Abstract: Brain field potentials are scale-free: their power spectra follow a $1/f^{\beta}$ law whose aperiodic exponent $\beta$ tracks cortical state, and sleep depth in particular is a shift in $\beta$.

By Dibakar Sigdel
llmsbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 4

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

arXiv:2608. 01624v1 Announce Type: cross Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a handful of scalars.

By Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim, Unggi Lee
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea