Hugging Face Trending Papers

Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures

Deep neural networks trained on natural images are shown to produce outputs consistent with human observers for brightness illusions. While this phenomenon has been documented across architectures, all evidence, to date, is measured at the output level: restored pixels, decoded trajectories, or classification decisions.

Hugging Face Trending Papers
Jun 3

Coarse-to-fine Hierarchical Architecture with Sequential Mamba for Brain Reconstruction

Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.

arXiv AI
Aug 28

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

The paper introduces Illusion-Reasoning, a benchmark that uses real-world visual illusion images paired with annotated question-answer pairs to evaluate Large Vision Language Models (LVLMs). It argues that current evaluations focus too narrowly on perception or specific domains, and that a joint assessment of perception and reasoning in open-world settings is lacking. Using this benchmark, the authors demonstrate that many LVLMs’ reasoning abilities are less advanced than previously claimed, offering new insights and directions for optimization.

By Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li
arXiv Machine Learning
Aug 4

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

arXiv:2608. 01821v1 Announce Type: cross Abstract: Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost.

By Yongkang Zhou, Xiang Xia, Cheng Yan, Fan Xu, Wuyang Zhang
arXiv AI
Jun 30

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

arXiv:2606. 29431v1 Announce Type: new Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image.

By Yichen Guo, Kai Tang, Fenglai Lin, Yiding Sun, Dongshuo Zhang, Wenya Wang, Lin William Cong, Shanghang Zhang
arXiv AI
Jul 8

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.

By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia