arXiv AI

ECI: Effective Contrastive Information to Evaluate Hard-Negatives

arXiv:2603. 20990v2 Announce Type: replace-cross Abstract: Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation.

arXiv Machine Learning
Jun 2

When Hard Negatives Hurt: Bridging the Generative-Discriminative Gap in Hard Negative Synthesis for Retrieval

arXiv:2606. 01304v1 Announce Type: new Abstract: Hard negative mining has become the dominant strategy for training retrievers, yet it faces intrinsic limitations: negatives are bounded by corpus availability, selected by retriever score rather than diagnostic value, and increasingly contaminated by false positives as the retriever improves.

By Zhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li, Jieming Zhu, Jingyu Li, Qianhui Zhu, Fengyuan Lu, Wang Jiaheng, Gang Wang, Hai-Tao Zheng, Zhaocheng Du
arXiv AI
Aug 20

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

The paper introduces HN-CLIP, a new objective for dense-caption retrieval that adapts similarity margins per negative example using the text encoder’s own geometry. By adding a detached caption‑similarity matrix to the negative logits, HN‑CLIP addresses the issue of near‑duplicate captions that cause premature loss saturation in InfoNCE training. Experiments on four benchmarks show that HN‑CLIP outperforms leading methods by 2.4–4.3 R@1, trains 2.4× faster than GOAL and 5.4× faster than StructXLIP, and achieves the best full‑data baseline with only 20% of the training data.

By Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang
arXiv Machine Learning
Jun 30

The Voronoi Bottleneck: Capacity-Aware Dense Retrieval for Product Search

arXiv:2606. 28359v1 Announce Type: cross Abstract: Dense embedding retrieval compresses all relevance information into a single inner product, imposing a fundamental geometric limit -- the Voronoi Bottleneck -- on the number of query-document relevance patterns expressible at fixed embedding dimension (d).

By Charith Chandra Sai Balne, Rithwik Maramraju, Siddharth Pratap Singh, Rohit Upadhyay, Aditya Singh, Chittaranjan Tripathy, Yogananda Domlur Seetharama
arXiv AI
Jul 13

Contrastive Weak-to-strong Generalization

arXiv:2510. 07884v2 Announce Type: replace-cross Abstract: Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling.

By Houcheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang, Chen Gao, Xiang Wang, Xiangnan He, Yang Deng
arXiv AI
Aug 19

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.

By Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang