arXiv Machine Learning

Beyond the Simplex: Balanced Prototype Geometry for Scorer-Agnostic Open-Set Recognition

arXiv:2606. 01883v1 Announce Type: new Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical imaging.

arXiv Machine Learning
Aug 28

Common Geodesics Do Not Guarantee Fisher Consistency of the Structured SVM: Minimal Counterexamples and a Tree-Metric Classification

The paper demonstrates that the common-geodesic condition—where every output triple shares a geodesic point in a metric—does not ensure Fisher consistency for the structured SVM with the standard coordinate-wise argmax decoder. It presents minimal counterexamples, including a four-output star and a tree-metric classification, showing that only path-shaped trees maintain argmax consistency. The study also identifies the smallest full-support counterexamples and provides exact primal-dual certificates for all optimality claims.

By Jintao Fei, Jiangying Luo
arXiv Machine Learning
Aug 27

Common-Center Geometry and Certified Radial Reconstruction for Energy-Form Full Conformal Regions

The paper investigates the geometry of full conformal prediction (FullCP) regions produced by an empirical energy‑form pairwise score. It shows that convexity of the candidate score alone does not ensure connected FullCP regions, and establishes conditions under which comparison regions share a common minimizer, making the exact conformal region star‑shaped. For power distances with exponent β≥1 the geometry is deterministic, and for β between 1 and 2 explicit Lipschitz bounds allow certified inner and outer radial envelopes with Hausdorff guarantees.

By Yiheng Feng
Hugging Face Trending Papers
Jul 30

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.

arXiv Machine Learning
Sep 10

Sharp Structure-Agnostic Minimax Risk for Partial Linear Models

arXiv:2609. 07997v1 Announce Type: new Abstract: We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025).

By Haichen Hu, David Simchi-Levi
arXiv AI
Sep 4

Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data

DIFFINT is a reconstruction‑based anomaly detector that uses a differentiable autoencoder with a latent bottleneck composed of soft, axis‑aligned interval memberships. Each latent unit represents a human‑readable hyper‑rectangle in feature space, allowing the model to encode how strongly an instance falls inside each interval and to compute reconstruction error as the anomaly score. The method provides a certified lower bound on reconstruction error for points outside all active intervals, a suppression mechanism for sparse abnormalities, and a closed‑form, label‑free importance ranking for each (unit, feature) pair, achieving top performance on 48 ADBench benchmarks against 22 baselines.

By Lamine Diop, Marc Plantevit
arXiv Machine Learning
Aug 18

The Limits of Binding in Dual Encoders

arXiv:2608. 15971v1 Announce Type: new Abstract: Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance when asked to distinguish "a red car and a blue dog" from "a blue car and a red dog".

By Kin Ian Lo