Hugging Face Trending Papers
Aug 4

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations.

arXiv Machine Learning
Sep 16

Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

The paper studies how to allocate a fixed computational budget between pretraining and fine‑tuning in a two‑stage ridge regression setting. By modeling the process as a compute‑split problem and analyzing data‑dependent evaluation geometries, it derives the optimal split in terms of prediction‑relevant spectral components of the pretraining and fine‑tuning empirical covariances. The analysis uses a basis‑invariant eigenspace decomposition and perturbative control of non‑commuting dynamics to capture how pretraining directions influence downstream predictions.

By Alex Buna, Fanghui Liu, Patrick Rebeschini
arXiv Machine Learning
Jul 31

Towards Stability of Parameter-Free Optimization

arXiv:2405. 04376v4 Announce Type: replace Abstract: Hyperparameter tuning, particularly the selection of an appropriate learning rate in adaptive gradient training methods, remains a challenge.

By Yijiang Pang, Shuyang Yu, Bao Hoang, Jiayu Zhou