arXiv AI

High-Dimensional Concentration and Retrieval Instability in Embedding Spaces: Implications for Retrieval-Augmented Generation

arXiv:2606. 28330v1 Announce Type: cross Abstract: Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semantic relevance.

arXiv Computer Vision
Aug 31

Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems

The paper reviews 50 image augmentation and generation techniques, categorizing them into ten groups, and conducts a large‑scale empirical study to assess their effectiveness as test generators for embedding‑based image retrieval systems. Using Amazon Titan and OpenCLIP embeddings, the authors evaluate the techniques across four dimensions—embedding‑space similarity, embedding uncertainty, semantic realism, and retrieval failure rate—on CIFAR‑10, ImageNet‑1K, and an industrial dataset. Results show that weather simulation and SaSPA yield the highest uncertainty and failure rates while maintaining realistic visuals, whereas GAN‑based methods produce low realism due to synthetic artifacts.

By Yehan De Silva, Anirudh Sridhar, Armin Lotfy, Nafiseh Kahani, Yvan Labiche, Ziyu Wang, Frank Ouyang, Clare Carty, Azalia Shamsaei
arXiv AI
Sep 4

WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval

WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval proposes a new approach to address information asymmetry in cross-modal retrieval. The method introduces Adaptive Entropy Thresholding to calibrate uncertainty, Asymmetry-aware Wildcard Decoding to emit wildcards instead of forced identifiers, and Blind-Spot Re-ranking to evaluate expanded candidates. Experiments on the M-BEIR benchmark show that WIDE outperforms existing generative retrieval methods by reducing forced hallucination while keeping index structures compact.

By Teng Guo, Xin Wang, Jiayou Xu, Keying Zhou, Jifeng Shen, Haoxin Ruan