arXiv Computer Vision

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

arXiv AI
Sep 30

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

NeuronEye is a plug‑in framework that builds a sparse, concept‑level neuron vocabulary from intermediate vision‑language model (VLM) representations and selectively activates query‑relevant visual concepts during inference. It decomposes vision‑token states into an overcomplete sparse basis organized by concept clusters, uses the language query to activate relevant clusters, localizes the corresponding image patches, and injects the focused evidence back into the vision tokens, while a suppression mechanism attenuates dominant perceptual directions. Experiments on Qwen2.5‑VL‑7B and LLaVA‑1.6‑7B show that NeuronEye improves CV‑Bench overall accuracy by +3.1, boosts Distance by +9.5, and raises BLINK Multi‑view by +8.3, indicating that sparse neuron vocabularies can act as active interfaces for concept‑level visual reasoning.

By Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao
arXiv Computer Vision
Sep 24

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.

By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv AI
Sep 2

When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection

The paper introduces Inverted Contrastive Learning for Unsupervised Feature Selection (ICLFS), a method that treats each feature as a sample by inverting the data matrix and applies a contrastive learning framework to learn consistent representations across masked positive views and a shuffled negative view. Feature saliency is derived from the magnitude of projector‑space embeddings, and a Laplacian‑Gated Ranking Correction step refines the ranking by reducing local redundancy. Experiments on 12 benchmark datasets show that ICLFS achieves the best clustering accuracy on 10 datasets compared to both classical and neural baselines, demonstrating the effectiveness of feature‑wise contrastive consistency for unsupervised feature selection.

By Utsab Ghosh, Roshni Chakraborty