← Back to all news
arXiv Machine Learning October 7, 2026 By Martin Eppert, Krishna Balasubramanian, Subhro Ghosh, Jason Klusowski, Yan Shuo Tan

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 7

Contaminated Multi-task Learning with Heterogeneity: Fundamental Limits and Optimal Algorithms

arXiv:2607. 02681v1 Announce Type: cross Abstract: Integrating information across related tasks can improve estimation and prediction in transfer, multi-task, and federated learning, but contamination and heterogeneity make robust borrowing challenging.

By Ye Tian, Mengchu Li, Marco Avella Medina
benchmarks
More like this →
arXiv Machine Learning
Aug 11

Training-Free Universal Approximation by Prompting Random Transformers

arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?

By Alexander Hsu, Rongjie Lai
llms
More like this →
arXiv Machine Learning
Jul 7

Gradient Descent as Implicit EM in Distance-Based Neural Models

arXiv:2512. 24780v2 Announce Type: replace Abstract: Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking.

By Alan Oursland
llms
More like this →
arXiv Machine Learning
Jul 22

Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention

arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.

By Subham Singh, Ashutosh Mishra, Subha Raut
llmssafety
More like this →
arXiv AI
Sep 28

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.

By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
llms
More like this →
arXiv Machine Learning
Jun 17

Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective

arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.

By Etienne Boursier, Claire Boyer
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea