Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Read the original on arXiv AI →The paper introduces HN-CLIP, a new objective for dense-caption retrieval that adapts similarity margins per negative example using the text encoder’s own geometry. By adding a detached caption‑similarity matrix to the negative logits, HN‑CLIP addresses the issue of near‑duplicate captions that cause premature loss saturation in InfoNCE training. Experiments on four benchmarks show that HN‑CLIP outperforms leading methods by 2.4–4.3 R@1, trains 2.4× faster than GOAL and 5.4× faster than StructXLIP, and achieves the best full‑data baseline with only 20% of the training data.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.