arXiv Machine Learning

Localizing Transfer Between Memorization Tasks

The paper investigates how pre‑training on random input‑output mappings (memorization tasks) can transfer to downstream tasks. It discovers two unexpected patterns: equivalent transfer, where each pre‑training epoch saves roughly one fine‑tuning epoch, and non‑equivalent transfer, where pre‑training on a mismatched task can be more efficient than training directly on the downstream task. Ablation studies reveal that transfer consists of a trivial magnitude‑driven effect in the last layer and a non‑trivial structure‑driven effect linked to covariance in other layers.

arXiv AI
Sep 1

Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

The study investigates how post‑training of large autoregressive language models (ARMs) into masked diffusion models (MDMs) affects their internal computation. Across two 7B ARM‑MDM families and four diagnostic tasks, the authors find that MDMs retain much of the ARM’s high‑attribution pathways on prefix‑dominant tasks, but reorganize computation toward earlier layers on globally constrained tasks. Component‑level probes reveal that ARMs depend on sharply specialized components, whereas MDMs show weaker specialization and more diffuse output‑space alignment.

By Injin Kong, Hyoungjoon Lee, Yohan Jo
Hugging Face Trending Papers
5d ago

Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases

The paper introduces ATLAS, a method that uses retained-domain activation atlases to guide task-specific fine‑tuning of language models. By providing local reference centers and directional filters, ATLAS learns a low‑rank residual that adapts to new tasks while preserving existing behavior. Experiments on Qwen3‑8B and other backbones show lower retained‑output KL divergence, fewer rewritten answers, and more stable responses compared to seven baselines.

arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
arXiv Machine Learning
Jun 4

Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency Transform

arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.

By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
arXiv AI
Sep 18

Why $\beta_1 = \beta_2$ Is Dynamically Special in Adam

The paper investigates why setting the two momentum parameters of Adam equal (β1=β2) has a special dynamic effect. By analysing Adam in continuous time, the authors show that the update decomposes into a sign component, a magnitude‑lag term proportional to the difference between the two memory times, and other terms. This lag term disappears exactly when β1=β2, making the diagonal the only regime where the mismatch‑induced response is structurally absent. Experiments on six vision and language tasks confirm that tied configurations are sign‑dominated, have smaller lag contributions, and exhibit smoother update‑norm trajectories.

By Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ort\'i