The Visual Target Matters: Learning across the Visual Hierarchy for Brain-to-Image Retrieval
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.24136v1 Announce Type: new Abstract: Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented...
arXiv:2606. 00121v1 Announce Type: cross Abstract: Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task in brain decoding.
Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.
The paper presents a reproducible single‑subject baseline for reconstructing visual stimuli from EEG using a temporal‑spatial convolutional encoder that maps averaged EEG signals to 512‑dimensional ViT-B/32 image features. On the THINGS‑EEG2 dataset, the model achieves 12.83%, 39.17%, and 58.00% image recall at ranks 1, 5, and 10, respectively, outperforming analytical chance levels. The study also shows that performance drops sharply when applying a model trained on one subject to others, and that direct conditional generators without external visual weights produce noise‑dominated outputs, indicating that only coarse semantic decoding is feasible under the tested protocol.
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
The paper introduces NEAR, a neural-anchor-based retrieval framework that improves brain-to-image retrieval when only a few neural trials are available. By treating a high‑repetition center as an anchor, NEAR uses a denoiser to pull noisy queries toward the anchor and a small network to predict pseudo anchors for candidate images, thereby aligning both neural and visual representations. Experiments on EEG, MEG, and fMRI datasets show consistent gains, including a 5.7–9.3 percentage point improvement in 200‑way Top‑1 accuracy on THINGS‑EEG2 with only one or four repetitions.