arXiv:2606.03715v3 Announce Type: replace
Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that co...
By Nurit Spingarn, Noa Cohen, Tamar Rott Shaham, Tomer Michaeli
arXiv:2607. 03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs).
By Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
The paper investigates how vision‑language models (VLMs) perform optical character recognition (OCR) by identifying attention heads that are causally necessary for OCR across four models. These heads are shown to be general‑purpose, producing interpretable semantic features for any image token, such as recognizing the word "bike" or the concept "feathers". By collapsing the heads’ attention weights into a verbalization lens transformation, the authors reveal that image representations align with language from early layers and can even be used to edit non‑word concepts in images, demonstrating the broader utility of this subspace.
By Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau
arXiv:2606. 03093v1 Announce Type: new Abstract: Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representations to produce behavior.
By Fan L. Cheng, Nikolaus Kriegeskorte
Steering Fields introduce adaptive vector fields that re-estimate steering directions at each step of a flow-based text-to-image generation process, replacing the fixed global steering vectors traditionally used. By operating on noisy states, they provide a continuous trade-off between steering strength and content preservation, and allow simultaneous induction and inhibition of concepts without explicit spatial masks or object priors. The method achieves state-of-the-art safety steering benchmarks and can also function as a structure-preserving image-editing technique, delivering high semantic fidelity while remaining model-agnostic and inversion-free.
By Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
By Tuc Nguyen, Thai Le
arXiv:2606. 09287v1 Announce Type: new Abstract: Understanding how transformer representations evolve across layers, not merely what they encode, remains an open problem in mechanistic interpretability.
By Vishal Pandey, Gopal Singh
arXiv:2607. 16295v1 Announce Type: cross Abstract: Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm.
By Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu
arXiv:2609.05575v1 Announce Type: new
Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability,...
By Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, Dianbo Liu
arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.
By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia