Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606.03715v3 Announce Type: replace Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that co...
arXiv:2607. 03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs).
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
The paper investigates how vision‑language models (VLMs) perform optical character recognition (OCR) by identifying attention heads that are causally necessary for OCR across four models. These heads are shown to be general‑purpose, producing interpretable semantic features for any image token, such as recognizing the word "bike" or the concept "feathers". By collapsing the heads’ attention weights into a verbalization lens transformation, the authors reveal that image representations align with language from early layers and can even be used to edit non‑word concepts in images, demonstrating the broader utility of this subspace.
arXiv:2606. 03093v1 Announce Type: new Abstract: Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representations to produce behavior.