Learning to Predict Distributions over Weight Updates for Test-Time Adaptation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces query‑conditioned hypernetworks that predict distributions over LoRA weight updates for large language models. By learning a distribution rather than a single point estimate, the method allows sampling multiple adapted models for the same query, improving performance over deterministic hypernetworks and token‑sampling baselines. The study also shows that these learned updates can transfer across different queries, indicating reusable adaptation patterns.
arXiv:2607. 19604v1 Announce Type: cross Abstract: Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge.
arXiv:2604. 26170v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge.
arXiv:2602. 05988v2 Announce Type: replace Abstract: Pre-training Large Language Models (LLMs) on web-scale datasets becomes fundamental for advancing general-purpose AI.
arXiv:2510. 16077v2 Announce Type: replace-cross Abstract: Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains without catastrophic forgetting.
arXiv:2609.37687v1 Announce Type: new Abstract: Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively queryin...