The paper introduces NEAR, a neural-anchor-based retrieval framework that improves brain-to-image retrieval when only a few neural trials are available. By treating a high‑repetition center as an anchor, NEAR uses a denoiser to pull noisy queries toward the anchor and a small network to predict pseudo anchors for candidate images, thereby aligning both neural and visual representations. Experiments on EEG, MEG, and fMRI datasets show consistent gains, including a 5.7–9.3 percentage point improvement in 200‑way Top‑1 accuracy on THINGS‑EEG2 with only one or four repetitions.
By Zhenyao Cui, Siyuan Kan, Dingkun Liu, Dongrui Wu
SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval proposes a new framework that aligns EEG signals from different subjects into a common image space without requiring labeled calibration data. By training on source subjects and estimating an orthogonal transformation at deployment, SCORE recovers target EEG coordinates and selects reliable EEG-image landmarks through hubness-corrected matching. The method achieves state‑of‑the‑art Top‑1/Top‑5 accuracy on two public benchmarks, outperforming existing baselines by significant margins.
By Zhenyao Cui, Siyuan Kan, Siyang Li, Ziwei Wang, Dongrui Wu
arXiv:2609.24136v1 Announce Type: new
Abstract: Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented...
By Ye Wang, HaoKun Ren, Hong Yu, Ruirui Li, Xiao Li, Ke Liu, Wei Wu
The paper introduces an adaptive cortically constrained method for aligning EEG signals with visual representations in zero‑shot brain‑to‑image retrieval. It reconstructs EEG into ROI‑level source patterns, encodes them with a Neuro‑ROI Attention Encoder, and applies evidence‑based adaptive visual supervision to account for response‑wise variability. On the THINGS‑EEG dataset, the approach achieves strong 200‑way retrieval performance and offers ROI‑level attribution for interpretability.
By Ye Wang, Haokun Ren, Wei Wu, Guoyin Wang, Zhuliang Yu, Hong Yu, Ke Liu
Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented by pretrained vision models, whose internal repr...
ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding introduces a model‑agnostic framework that adaptively aligns EEG signals with visual semantics. It replaces fixed visual or textual anchors with EEG‑aware class‑level contrastive supervision and employs structure‑consistent interpolation to preserve channel‑wise and temporal importance. Across multiple evaluation settings—including subject‑dependent, subject‑independent, strict cross‑subject transfer, and continual adaptation—ProCA delivers significant performance gains, achieving relative Top‑1 improvements ranging from 7.4% to 28.1%.
By Kanglei Zhou, Chunyan Lan, Dongyang Li, Jun Zhu, Liyuan Wang
The paper introduces CERES, a closed‑loop multimodal indexing framework that addresses semantic collapse in multimodal generation by building a three‑level semantic pyramid and using scale‑routed cross‑attention to generate images that remain retrievable by their original queries. CERES employs a co‑occurrence‑aware router, a lightweight U‑Net generator, and a soft‑Jaccard coverage objective to ensure generated images cover the intended concepts, verified by re‑indexing with a frozen vision‑language model and an external DINOv2 probe. Experiments on four pansharpening benchmarks show state‑of‑the‑art performance, especially under extreme scale variation, and significant improvements in concept‑query retrieval and image‑text ranking metrics.
By Guangyuan Dong, Chuang Liu, Haoyu Wang, Yangchen Zeng, Jiaqi Zhang, Li Jiuxing, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin, Alexander Lim Han Yang, Yusen Wu
arXiv:2608. 20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve.
By Guangyuan Dong, Chuang Liu, Yangchen Zeng, Haoyu Wang, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin
The paper introduces TFA, a training‑free aggregation technique that calibrates frozen visual foundation models for visual place recognition. TFA uses cross‑codebook agreement, retrieval coverage, and spectral statistics to adjust residual assignment, spectral shaping, and global‑feature fusion without requiring place labels or task‑specific weights. Experiments with a DINOv2‑B backbone show significant Recall@1 gains over existing training‑free methods across multiple benchmarks, demonstrating that reliability‑guided aggregation can unlock additional retrieval performance from frozen representations.
By Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang
arXiv:2607. 12364v1 Announce Type: cross Abstract: EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning.
By Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
The paper presents a reproducible single‑subject baseline for reconstructing visual stimuli from EEG using a temporal‑spatial convolutional encoder that maps averaged EEG signals to 512‑dimensional ViT-B/32 image features. On the THINGS‑EEG2 dataset, the model achieves 12.83%, 39.17%, and 58.00% image recall at ranks 1, 5, and 10, respectively, outperforming analytical chance levels. The study also shows that performance drops sharply when applying a model trained on one subject to others, and that direct conditional generators without external visual weights produce noise‑dominated outputs, indicating that only coarse semantic decoding is feasible under the tested protocol.
By Harshit Goyal
arXiv:2606. 15782v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation.
By Pratheswaran Hariharan, Haiping Xu, Donghui Yan