arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.
By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu
arXiv:2601.21948v2 Announce Type: replace
Abstract: Neural visual decoding is a central problem in brain-computer interface research, aiming to reconstruct human visual perception and to elucidate th...
By Yang Du, Siyuan Dai, Yonghao Song, Paul M. Thompson, Haoteng Tang, Liang Zhan
arXiv:2606. 30319v1 Announce Type: cross Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience.
By Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng Song, Changqing Zhang, Jiamin Wu
arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.
By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv:2607. 12364v1 Announce Type: cross Abstract: EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning.
By Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
arXiv:2606. 00121v1 Announce Type: cross Abstract: Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task in brain decoding.
By Yizhuo Lu, Changde Du, Qiongyi Zhou, Liuyun Jiang, Huiguang He
arXiv:2607. 15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI.
By Bo-An Chang, Yu-Chih Chen
The paper introduces MD‑SigLIP, a margin‑regularized structured semantic alignment framework that directly aligns brain embeddings with text embeddings in a shared semantic space for retrieval‑based decoding. It builds on duplicate‑aware sigmoid contrastive learning and adds a listwise margin‑regularized term to enforce structured ranking constraints between positive semantic clusters and negative samples. Experiments show that this approach achieves state‑of‑the‑art retrieval performance in both full‑vocabulary and subset evaluation settings.
By Jiaqi Wang, Huawen Hu, Shu Zhang
SVG-Score introduces a human‑aligned evaluation framework for text‑to‑SVG generation, addressing the shortcomings of existing image‑based metrics like CLIPScore that poorly capture SVG‑specific errors such as color, count, and spatial inaccuracies. The authors first demonstrate that CLIP‑based scores are largely insensitive to these errors and that generic Vision‑Language Models respond inconsistently across error types and styles. They then present a human‑annotated Semantic Alignment dataset and develop two complementary evaluators: a CLIP‑based scorer adapted to vector graphics and a VLM judge refined through supervised fine‑tuning and reward‑shaped reinforcement learning, enabling both fast large‑scale and expressive, interpretable assessment of SVG generators.
By Marco Cipriano, Leonardo Zini, Alexandra Schild, Valentin Teutschbein, Afsana Mimi, Marcella Cornia, Lorenzo Baraldi, Gerard de Melo
arXiv:2606. 16799v1 Announce Type: cross Abstract: Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations.
By Zijie Meng
arXiv:2607. 18237v1 Announce Type: cross Abstract: Human visual similarity judgments are context-dependent.
By Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang