The paper introduces an unsupervised framework that merges manifold learning with rank‑based interpretable graph embeddings to address the Geometric and Interpretability Gaps in visual representation learning. By first analyzing contextual information on the dataset manifold and then producing sparse, self‑explainable embeddings, the method achieves dimensionality reduction while preserving or improving performance in image retrieval and semi‑supervised Graph Convolutional Network classification. Experiments across varied datasets confirm that these context‑aware representations maintain high downstream effectiveness.
By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Vinicius Atsushi Sato Kawai, Daniel Carlos Guimar\~aes Pedronette
arXiv:2608.29001v1 Announce Type: new
Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these conne...
By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Lucas Pascotti Valem, Andr\'e Freitas, Daniel Carlos Guimar\~aes Pedronette
CORE improves compositional reasoning in multimodal language models by distilling a cross‑attentive reranker’s fine‑grained judgments into the embedding model. It generates candidate lists across five compositional matching levels and trains with a Rank‑KL objective to replicate the reranker’s ranking. Experiments on COLA, SUGARCREPE++, and NEGBENCH show CORE‑RERANKER‑8B outperforms Jina‑Reranker by 10.7 points, while CORE‑EMBED‑8B achieves the best overall average among evaluated embeddings, with gains also transferring to the MCMR benchmark without harming COCO or Flickr30K retrieval.
By Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu
The paper introduces SRAIN, a framework that learns sample‑wise, rank‑aware interpolation weights for composed visual data retrieval. Instead of relying on complex multimodal large language models, SRAIN uses simple linear interpolation in embedding space, dynamically predicting query‑specific weights through batch‑wise rank‑aware estimation and a compact memory bank for hard negatives. This approach achieves state‑of‑the‑art performance on composed video retrieval and competitive results on composed image retrieval while significantly reducing query‑time latency.
By Boseung Jeong, Taegyu Park, Donghyeon Kwon, Hyunsouk Cho, Suha Kwak
arXiv:2606. 17406v1 Announce Type: cross Abstract: Feature extraction involves the identification and extraction of salient characteristics or patterns, including edges, textures, shapes, and color attributes.
By Marina Chagas Bulach Gapski, Vinicius Atsushi Sato Kawai, Gustavo Rosseto Leticio, Lucas Pascotti Valem, Daniel Carlos Guimar\~aes Pedronette, Mohand Said Allili
arXiv:2606. 04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional space.
By Mohammad Tariqul Islam, Jason W. Fleischer
The paper introduces Inverted Contrastive Learning for Unsupervised Feature Selection (ICLFS), a method that treats each feature as a sample by inverting the data matrix and applies a contrastive learning framework to learn consistent representations across masked positive views and a shuffled negative view. Feature saliency is derived from the magnitude of projector‑space embeddings, and a Laplacian‑Gated Ranking Correction step refines the ranking by reducing local redundancy. Experiments on 12 benchmark datasets show that ICLFS achieves the best clustering accuracy on 10 datasets compared to both classical and neural baselines, demonstrating the effectiveness of feature‑wise contrastive consistency for unsupervised feature selection.
By Utsab Ghosh, Roshni Chakraborty
arXiv:2608. 12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly.
By Kaipeng Li, Haitao Yu, Xuanchen Zhou
arXiv:2608. 11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning.
By Archan Dutta, Vyanktesh Kanungo
arXiv:2601. 20844v3 Announce Type: replace-cross Abstract: This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of $m$ object vectors so that every subset of size at most $k$ is exactly retrieved by score comparison.
By Zihao Wang, Hang Yin, Lihui Liu, Hanghang Tong, Yangqiu Song, Ginny Wong, Simon See
arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.
By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
arXiv:2608. 20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve.
By Guangyuan Dong, Chuang Liu, Yangchen Zeng, Haoyu Wang, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin