arXiv Computer Vision By Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu

Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization

Read the original on arXiv Computer Vision →

The paper investigates how affective fine‑tuning shapes the internal architecture of multimodal foundation models. By analyzing 13 model instances across nine designs, it finds that adapting the feed‑forward network (FFN) consistently outperforms attention‑only adaptation and nearly matches full‑model tuning, revealing the FFN as an efficient adaptation substrate. Moreover, joint optimization leads to emergent functional specialization, notably a prominent gate projection pathway, which the authors exploit in Gate‑Focused Efficient Tuning (GET) to achieve 96.2–98.0% of full‑model performance with only 19.3–24.5% of the trainable parameters.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
arXiv AI
Jun 2

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

arXiv:2606. 02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (dense transformer, mixture-of-experts) and two pretraining corpora (The Pile, DCLM): Pythia 1B, OLMo 1B-0724-hf, and OLMoE 1B-7B-0924.

By Yongzhong Xu