← Back to all news
arXiv Machine Learning September 30, 2026 By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State

Predictive Geometry of Hidden Trajectories in Transformers

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 14

Geometric and Behavioral Stratification in Transformer Residual Streams

arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.

By Nelson Guda
llms
More like this →
arXiv Machine Learning
Sep 15

Disentangling Representation Evolution in Transformers through Directional Decomposition

arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...

By Shwai He, Haichao Zhang, Shen Yan
llmsrobotics
More like this →
arXiv Machine Learning
Sep 24

Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces

arXiv:2609.27988v1 Announce Type: cross Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...

By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
llmscomputer-visionfine-tuningefficiency
More like this →
Hugging Face Trending Papers
Sep 23

Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces

Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...

llmscomputer-visionfine-tuningefficiency
More like this →
arXiv AI
3d ago

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.

By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
llms
More like this →
arXiv Machine Learning
Jul 21

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

arXiv:2607. 17696v1 Announce Type: cross Abstract: We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions.

By Cheng Huan, Hongwei Yuan
llmssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea