arXiv AI

Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval

The paper introduces GeoPatch, a fixed‑scaffold patch encoder that decouples token support from signal geometry for time‑warp robust sequence retrieval. By using geometry‑derived descriptors (slope, curvature, acceleration, affine‑residual, confidence) as continuous conditioning variables, GeoPatch prevents patch boundary drift while still allowing geometry to modulate embeddings. Experiments on ECG, speech, and multivariate time‑series data show that GeoPatch improves early‑rank retrieval under timing variation and highlights a trade‑off between local surface matching and strict non‑overlap retrieval.

arXiv AI
Sep 24

TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval

The paper introduces TEMPS, a modular temporal branch that enhances semantic retrievers by adding a temporal scoring component. TEMPS resolves anchored temporal expressions into intervals, matches them to Gaussian distributions, and trains an anchor-date-conditioned encoder using grounding supervision without hand‑labeled data. On three temporal benchmarks, TEMPS improves MRR across all tested backbones and raises R@1 from 19.92 to 25.39 on the TS‑Retriever, surpassing prior temporal state‑of‑the‑art methods.

By Mourad Hassani, Julien Romero, Amel Bouzeghoub, Christian Jacquelinet
arXiv AI
Jul 16

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs

arXiv:2605. 16366v2 Announce Type: replace-cross Abstract: Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many spatial tokens, while capturing short-lived events requires dense temporal sampling.

By Yigui Feng (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Qinglin Wang (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Yang Liu (The Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, Guangdong, China), Jie Liu (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China)
arXiv Computer Vision
2d ago

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

PAGER is a label‑free adaptation method that aligns partial, viewpoint‑dependent 3D observations with a frozen global semantic space. It uses matched‑point feature alignment and relational supervision to anchor partial features to their global counterparts while preserving similarity structure, all without altering the pretrained encoder or global probe. Experiments show PAGER outperforms label‑supervised PEFT on Sonata and Concerto, and achieves superior zero‑shot transfer from ScanNet to ScanNet++ compared to fully fine‑tuned Sonata.

By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni