arXiv:2601. 09173v5 Announce Type: replace Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar.
By Prashant C. Raju
arXiv:2606. 00124v1 Announce Type: cross Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial representations is not well understood.
By Mahmoud Mannes
The study investigates how to automatically predict the build orientation for selective laser melting (SLM) of dental parts using supervised machine learning. Researchers trained two different neural network backbones—ResNet‑50 on multi‑view images and PointNeXt‑S on point clouds—on about 2,400 patient‑specific parts, evaluating 13 different ways to represent the up‑axis (six classical SO(3) parameterizations and seven unit‑sphere representations). They found that applying test‑time augmentation (TTA) over 21 known rotations consistently reduced angular error, with the octahedral map achieving the lowest mean error (10.6°) on ResNet‑50, while direct S² representations performed best overall but may be influenced by label noise.
By Felix Schmalzel, Reimar Waitz, Moritz Kronberger, Thorsten Sch\"oler
arXiv:2406. 07049v3 Announce Type: replace-cross Abstract: Understanding spatial relationships across all dimensions is fundamental for intelligent systems.
By Boyang Li, Yulin Wu, Nuoxian Huang, Wenjia Zhang
arXiv:2509. 11218v2 Announce Type: replace-cross Abstract: Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification.
By Johann Schmidt, Sebastian Stober
arXiv:2505. 21736v2 Announce Type: replace-cross Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision.
By Siqi Fang, Zachary Schlamowitz, Andrew Bennecke, Daniel J. Tward
arXiv:2608. 08173v1 Announce Type: cross Abstract: Longitudinal MRI enables sensitive measurement of structural brain change for studying aging and neurodegenerative disease.
By Jingru Fu, Kathleen E. Larson, Douglas N. Greve, Bruce Fischl, Malte Hoffmann
arXiv:2607. 26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume.
By Mahesh Godavarti
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
The study investigates how the geometry of representations in artificial neural networks can be steered to improve bidirectional alignment with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The changes also lowered effective dimensionality and reorganized the shared subspace, making forward and reverse predictivity more symmetric at certain spectral exponents.
By Samuel Kostousov, Abhinn Kaushik, Brokoslaw Laschowski
The study investigates how aligning the representational geometry of artificial neural networks can improve bidirectional predictivity with biological neural responses. By applying spectral regularization during training of self‑supervised contrastive vision models, the authors increased reverse predictivity by 55% while only modestly reducing forward predictivity. The adjustments also lowered effective dimensionality and reorganized the shared representational subspace, making forward and reverse predictivity more symmetric at intermediate spectral exponents.
arXiv:2606. 04922v1 Announce Type: cross Abstract: Current prompt-based and adapter-based tuning of vision-language models (VLMs) is attractive for medical imaging, where clinical data sensitivity favors frozen backbones and annotations are limited.
By Tran Dinh Tien, Zhiqiang Shen