arXiv:2602.15382v3 Announce Type: replace-cross
Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...
By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.
arXiv:2609.01200v1 Announce Type: new
Abstract: When the visual encoder and the language decoder of a vision-language model (VLM) run on different compute nodes, the intermediate visual-token embeddi...
By Reza Heidari, Hamed R. Tavakoli, Juho Kannala
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
NeuronEye is a plug‑in framework that builds a sparse, concept‑level neuron vocabulary from intermediate vision‑language model (VLM) representations and selectively activates query‑relevant visual concepts during inference. It decomposes vision‑token states into an overcomplete sparse basis organized by concept clusters, uses the language query to activate relevant clusters, localizes the corresponding image patches, and injects the focused evidence back into the vision tokens, while a suppression mechanism attenuates dominant perceptual directions. Experiments on Qwen2.5‑VL‑7B and LLaVA‑1.6‑7B show that NeuronEye improves CV‑Bench overall accuracy by +3.1, boosts Distance by +9.5, and raises BLINK Multi‑view by +8.3, indicating that sparse neuron vocabularies can act as active interfaces for concept‑level visual reasoning.
By Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao
arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.
By Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan