Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
The paper demonstrates that the common-geodesic condition—where every output triple shares a geodesic point in a metric—does not ensure Fisher consistency for the structured SVM with the standard coordinate-wise argmax decoder. It presents minimal counterexamples, including a four-output star and a tree-metric classification, showing that only path-shaped trees maintain argmax consistency. The study also identifies the smallest full-support counterexamples and provides exact primal-dual certificates for all optimality claims.
By Jintao Fei, Jiangying Luo
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2606. 17961v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arrangement of inputs.
By Andrea Santomauro, Luigi Portinale, Giorgio Leonardi
arXiv:2601. 20844v3 Announce Type: replace-cross Abstract: This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of $m$ object vectors so that every subset of size at most $k$ is exactly retrieved by score comparison.
By Zihao Wang, Hang Yin, Lihui Liu, Hanghang Tong, Yangqiu Song, Ginny Wong, Simon See
arXiv:2608.30254v1 Announce Type: new
Abstract: We resolve the threshold part of Question 4 of the COLT 2025 open problem "Data Selection for Regression Tasks" of Hanneke, Moran, Shlimovich and Yehud...
By Guangjian Zhang
The paper investigates the geometry of full conformal prediction (FullCP) regions produced by an empirical energy‑form pairwise score. It shows that convexity of the candidate score alone does not ensure connected FullCP regions, and establishes conditions under which comparison regions share a common minimizer, making the exact conformal region star‑shaped. For power distances with exponent β≥1 the geometry is deterministic, and for β between 1 and 2 explicit Lipschitz bounds allow certified inner and outer radial envelopes with Hausdorff guarantees.
By Yiheng Feng
We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.
arXiv:2609. 07997v1 Announce Type: new Abstract: We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025).
By Haichen Hu, David Simchi-Levi
DIFFINT is a reconstruction‑based anomaly detector that uses a differentiable autoencoder with a latent bottleneck composed of soft, axis‑aligned interval memberships. Each latent unit represents a human‑readable hyper‑rectangle in feature space, allowing the model to encode how strongly an instance falls inside each interval and to compute reconstruction error as the anomaly score. The method provides a certified lower bound on reconstruction error for points outside all active intervals, a suppression mechanism for sparse abnormalities, and a closed‑form, label‑free importance ranking for each (unit, feature) pair, achieving top performance on 48 ADBench benchmarks against 22 baselines.
By Lamine Diop, Marc Plantevit
arXiv:2608. 15971v1 Announce Type: new Abstract: Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance when asked to distinguish "a red car and a blue dog" from "a blue car and a red dog".
By Kin Ian Lo