arXiv Machine Learning

A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt

arXiv:2606. 15569v1 Announce Type: new Abstract: Test-time training (TTT) adapts a pretrained model to each prompt via parameter updates, improving accuracy under pretraining-to-test distribution shifts.

arXiv Machine Learning
Sep 11

Test time training enhances in-context learning of nonlinear functions

The paper studies how test‑time training (TTT) improves in‑context learning (ICL) for nonlinear models, focusing on single‑index models where features lie in a hidden low‑dimensional subspace. By applying TTT to single‑layer transformers trained with gradient‑based methods, the authors derive an upper bound on prediction risk and show that TTT allows the model to adapt to both feature vectors and link functions that vary across tasks—something ICL alone struggles to achieve. They also provide a convergence rate indicating that predictive error can approach the noise level as context size and network width increase.

By Kento Kuwataka, Taiji Suzuki
arXiv Machine Learning
Sep 11

Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees

The paper introduces a statistical framework for post‑training hyperparameter selection, emphasizing the learn‑then‑test (LTT) paradigm. It treats hyperparameter tuning as a multiple hypothesis testing problem over a candidate set, enabling the selection of hyperparameters that meet specified reliability constraints such as risk bounds or information‑theoretic limits. The framework provides finite‑sample control of error probabilities using p‑values, e‑values, and concentration inequalities derived from first principles.

By Amirmohammad Farzaneh, Osvaldo Simeone
arXiv Machine Learning
Sep 3

Rethinking the Teacher-Student Framework for Test-Time Adaptation

The paper investigates the teacher‑student framework used in Test‑Time Adaptation (TTA) and questions the common practice of updating the teacher via an exponential moving average of the student. The authors demonstrate that error accumulation still occurs, especially over longer sequences, and propose an intransigent teacher that remains fixed. This modification yields significant performance gains across multiple datasets, longer scenarios, and various architectures, including semantic segmentation, while also improving robustness to hyperparameter changes.

By Damian S\'ojka, Marc Masana, Bart{\l}omiej Twardowski, Sebastian Cygert
arXiv AI
Aug 26

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

The paper introduces KENDO, a unified framework that combines Ensemble Gaussian Processes with disagreement‑aware acquisition strategies to address hyperparameter selection in Bayesian optimization and active learning. By replacing costly hyperparameter sampling with a kernel ensemble and adaptive Bayesian weighting, KENDO‑BO and KENDO‑AL provide self‑correcting mechanisms tailored to their respective tasks. Experiments on synthetic and real‑world benchmarks show that KENDO‑BO matches or outperforms state‑of‑the‑art methods while cutting computational cost up to fivefold, and KENDO‑AL delivers better predictive calibration with up to 27‑times speedup compared to MCMC‑based baselines.

By Heng Zhang, Haotian Xiang, Qin Lu, Konstantinos D. Polyzos, Tara Javidi
arXiv AI
Jul 13

Self-Guided Test-Time Training for Long-Context LLMs

arXiv:2607. 09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs.

By Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu