The paper introduces Hyper^2, a dual‑space consistency framework that applies hyperbolic geometry consistently to both the loss and the encoder in point‑cloud completion tasks. By reusing the same arcosh(1+αd²) function as a positional bias in refinement attention and as the Chamfer loss, Hyper^2 achieves significant Chamfer error reductions—up to 22.9% on ShapeNet‑55 and 37.5% on unseen ShapeNet‑34—while adding only ~1.6% FLOPs. The authors demonstrate that geometric consistency across encoder and loss, rather than either component alone, is key to effective hyperbolic supervision, supported by two model‑agnostic indicators that peak only when both are hyperbolic.
By Guantian Zheng, Haiyang Xu, Tianyu Gao
The paper investigates why hierarchical image retrieval improves when using frozen DINOv2 features. It compares Euclidean and hyperbolic embeddings trained with taxonomy-distance regression or a taxonomy-aware supervised contrastive objective, finding that the choice of loss function (objective family) contributes more to hierarchy-aware performance than the geometry of the embedding space. Semantic alignment of the taxonomy also plays a significant role, while stronger negative curvature does not explain the gains.
By Ling Shi (Southeast University)
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
The paper introduces SurgHiBench, a hierarchy-aware evaluation suite for surgical video understanding that measures recognition, consistency, and severity across different granularity levels. It also presents HyperSurg, a hyperbolic model that enforces phase-step containment using entailment cones, evaluated on four datasets covering three procedure types. The study shows that models with identical accuracy can differ significantly in error severity, and that hyperbolic geometry improves predictions by aligning them with the procedural hierarchy.
By Ana Manzano Rodr\'iguez, Pascal Mettes, Marlies P. Schijven, Cees G. M. Snoek
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
Surgical procedures follow a phase-to-step hierarchy, yet the video-language models used to recognize them are evaluated with flat per-level metrics that ignore cross-level coherence and error structu...