arXiv:2606. 17961v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arrangement of inputs.
By Andrea Santomauro, Luigi Portinale, Giorgio Leonardi
Prototype-based networks provide inherently interpretable classification by linking predictions to learned exemplars, but their use in 3D point clouds and clinical surface-pair reasoning remains limited. We introduce ProtoPointNet, a prototype-based model for dental occlusion classification from registered upper--lower intraoral arch pairs.
The paper presents a new monocular spacecraft pose estimation model that achieves the lowest reported mean rotation errors on the SPEED+ lightbox and sunlamp test sets. By replacing smaller encoders with a large self‑supervised ViT foundation model (DINOv3) and scaling up to 840 M parameters, the authors improve accuracy from 300 M to 840 M parameters without saturation. The 840 M model also runs on a Jetson Orin NX 16 GB with 133.8 ms per crop and 32.0 W power draw, demonstrating embedded inference feasibility while training solely on synthetic data.
By John Church, Vazghen Nikolian
arXiv:2609.38578v1 Announce Type: cross
Abstract: Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art...
By Kia-J\"ung Yang, Fabian H. Sinz, Pawe{\l} A. Pierzchlewicz
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
By Mark Oskin
arXiv:2608.31045v1 Announce Type: new
Abstract: Rotational symmetry is one of the most important structural principles in machine learning on 3D data. In applications ranging from physics and materia...
By Peter Lippmann, Fred A. Hamprecht
GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.
By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson
arXiv:2609.15083v1 Announce Type: new
Abstract: Mixed-curvature representation learning seeks to capture rich geometric structures that cannot be adequately modeled by a single curvature regime. Exis...
By Xingrun Li, Yusuke Mukuta, Xin Yang, Yinyu Ye, Tatsuya Harada
arXiv:2601. 09173v5 Announce Type: replace Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar.
By Prashant C. Raju
arXiv:2601.13913v3 Announce Type: replace
Abstract: We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2...
By Pavlo Melnyk, Cuong Le, Urs Waldmann, Per-Erik Forss\'en, Bastian Wandt
The paper introduces a Geometric-to-Semantic Spherical Transfer Learning framework for labeling cortical sulci on brain surfaces. It first pre‑trains a spherical encoder on ~30,000 unlabeled UK Biobank subjects using only curvature and depth, then injects sulcal fundi lines as a soft‑initialized Topological Prior Injector to bridge the geometric‑semantic gap. Experiments show the method surpasses fully supervised baselines, achieving a mean Dice score of 0.77 and delivering the largest gains on variable and tertiary sulci.
By Saeb Tounsi, Jo\"el Chavas, Pietro Gori, Vincent Frouin, Denis Rivi\`ere, Jean-Fran\c{c}ois Mangin
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...