arXiv Machine Learning

Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation

Hugging Face Trending Papers
Aug 4

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations.

arXiv Machine Learning
Sep 16

Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

The paper studies how to allocate a fixed computational budget between pretraining and fine‑tuning in a two‑stage ridge regression setting. By modeling the process as a compute‑split problem and analyzing data‑dependent evaluation geometries, it derives the optimal split in terms of prediction‑relevant spectral components of the pretraining and fine‑tuning empirical covariances. The analysis uses a basis‑invariant eigenspace decomposition and perturbative control of non‑commuting dynamics to capture how pretraining directions influence downstream predictions.

By Alex Buna, Fanghui Liu, Patrick Rebeschini
arXiv Machine Learning
Jul 31

Towards Stability of Parameter-Free Optimization

arXiv:2405. 04376v4 Announce Type: replace Abstract: Hyperparameter tuning, particularly the selection of an appropriate learning rate in adaptive gradient training methods, remains a challenge.

By Yijiang Pang, Shuyang Yu, Bao Hoang, Jiayu Zhou
arXiv AI
Sep 18

Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification

The paper introduces JANUS, a post‑hoc weight rectification framework that enforces Parameter Space Orthogonality to prevent catastrophic forgetting when fine‑tuning foundation models. By projecting updates into the Jacobian Null Space and employing a Multi‑step Adaptive Rectification mechanism, JANUS dynamically verifies trust regions and adjusts step sizes. Additional techniques such as ghost projection, ghost orientation comparison, and sequence‑level SVD compression provide temporal and spatial efficiency, enabling JANUS to integrate seamlessly with various fine‑tuning methods and effectively mitigate the stability‑plasticity dilemma.

By Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang, Jingliang Duan, Keqiang Li, Shengbo Eben Li
arXiv Computer Vision
Aug 24

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

The paper surveys Continual Test-Time Adaptation (CTTA), a framework that adapts pretrained computer‑vision models to non‑stationary target distributions without source data or labeled targets, while avoiding catastrophic forgetting and error accumulation. It formally defines the CTTA problem, categorizes existing methods into optimization‑based, parameter‑efficient, and architecture‑based families, and reviews representative techniques and benchmarks across standard evaluation settings. The survey also outlines current limitations and proposes future research directions, such as adapting foundation models and black‑box systems.

By Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang, Shuaicheng Niu, Taki Hasan Rafi, Jihun Hamm, Marco Pedersoli, Jose Dolz, Yunhui Guo