Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization
Read the original on arXiv Computer Vision →The paper investigates how affective fine‑tuning shapes the internal architecture of multimodal foundation models. By analyzing 13 model instances across nine designs, it finds that adapting the feed‑forward network (FFN) consistently outperforms attention‑only adaptation and nearly matches full‑model tuning, revealing the FFN as an efficient adaptation substrate. Moreover, joint optimization leads to emergent functional specialization, notably a prominent gate projection pathway, which the authors exploit in Gate‑Focused Efficient Tuning (GET) to achieve 96.2–98.0% of full‑model performance with only 19.3–24.5% of the trainable parameters.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.