arXiv AI By Di Wu, Xiaohui Zhu

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

Read the original on arXiv AI →

arXiv:2608. 10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

arXiv:2602.15382v3 Announce Type: replace-cross Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...

By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
Hugging Face Trending Papers
Jul 27

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.

arXiv AI
4d ago

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

NeuronEye is a plug‑in framework that builds a sparse, concept‑level neuron vocabulary from intermediate vision‑language model (VLM) representations and selectively activates query‑relevant visual concepts during inference. It decomposes vision‑token states into an overcomplete sparse basis organized by concept clusters, uses the language query to activate relevant clusters, localizes the corresponding image patches, and injects the focused evidence back into the vision tokens, while a suppression mechanism attenuates dominant perceptual directions. Experiments on Qwen2.5‑VL‑7B and LLaVA‑1.6‑7B show that NeuronEye improves CV‑Bench overall accuracy by +3.1, boosts Distance by +9.5, and raises BLINK Multi‑view by +8.3, indicating that sparse neuron vocabularies can act as active interfaces for concept‑level visual reasoning.

By Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao