arXiv Computer Vision By Katsuya Ogata, Yuta Nakashima

Virtual Encoders in Multimodal Transformers

Read the original on arXiv Computer Vision →

The paper investigates how multimodal language models can generate perceptual representations without dedicated encoders. It shows that the shared transformer can internally create these representations in its early-to-middle layers, a structure termed a Virtual Encoder. Experiments with linear probing, similarity metrics, and causal analysis reveal that this encoder-like computation emerges even when models receive only perceptual tokens, indicating that perception and language processing can be decoupled within a single architecture.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 11

LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

arXiv:2503. 06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging.

By Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer