arXiv Machine Learning By Tomoya Wakayama

A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt

Read the original on arXiv Machine Learning →

arXiv:2606. 15569v1 Announce Type: new Abstract: Test-time training (TTT) adapts a pretrained model to each prompt via parameter updates, improving accuracy under pretraining-to-test distribution shifts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Test time training enhances in-context learning of nonlinear functions

The paper studies how test‑time training (TTT) improves in‑context learning (ICL) for nonlinear models, focusing on single‑index models where features lie in a hidden low‑dimensional subspace. By applying TTT to single‑layer transformers trained with gradient‑based methods, the authors derive an upper bound on prediction risk and show that TTT allows the model to adapt to both feature vectors and link functions that vary across tasks—something ICL alone struggles to achieve. They also provide a convergence rate indicating that predictive error can approach the noise level as context size and network width increase.

By Kento Kuwataka, Taiji Suzuki
arXiv Machine Learning
Sep 11

Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees

The paper introduces a statistical framework for post‑training hyperparameter selection, emphasizing the learn‑then‑test (LTT) paradigm. It treats hyperparameter tuning as a multiple hypothesis testing problem over a candidate set, enabling the selection of hyperparameters that meet specified reliability constraints such as risk bounds or information‑theoretic limits. The framework provides finite‑sample control of error probabilities using p‑values, e‑values, and concentration inequalities derived from first principles.

By Amirmohammad Farzaneh, Osvaldo Simeone