arXiv AI

Temporal Preference Optimization for Unsupervised Retrieval

arXiv:2606. 17664v1 Announce Type: cross Abstract: Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled documents via contrastive learning, but they struggle to capture the temporal relevance, retrieving semantically related but temporally misaligned documents-an important aspect when a document collection spans multiple time periods (e.

arXiv AI
Sep 24

TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval

The paper introduces TEMPS, a modular temporal branch that enhances semantic retrievers by adding a temporal scoring component. TEMPS resolves anchored temporal expressions into intervals, matches them to Gaussian distributions, and trains an anchor-date-conditioned encoder using grounding supervision without hand‑labeled data. On three temporal benchmarks, TEMPS improves MRR across all tested backbones and raises R@1 from 19.92 to 25.39 on the TS‑Retriever, surpassing prior temporal state‑of‑the‑art methods.

By Mourad Hassani, Julien Romero, Amel Bouzeghoub, Christian Jacquelinet
arXiv AI
Aug 19

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.

By Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
arXiv Machine Learning
Sep 17

Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision

The paper investigates which historical examples are most useful for time‑series forecasting by defining predictive relevance as the expected future utility conditioned on inference‑time information. It introduces a two‑stage approach: a normalized‑pattern retriever generates a coarse candidate set, and a lightweight MLP reranks these candidates using future‑supervised relevance while keeping inference strictly past‑only. Experiments on six benchmarks show that this reranker improves pattern retrieval and outperforms a matched‑protocol baseline, revealing that historical relevance is structured, domain‑dependent, and not governed by a single universal retrieval rule.

By Yong-Hoon Choi, Kwang-Hyun Park, Youngjin Cho
arXiv Machine Learning
Jun 15

Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA

arXiv:2604. 23336v3 Announce Type: replace-cross Abstract: Unlike traditional fact-based retrieval, rationale-based retrieval typically necessitates cross-encoding of query-document pairs using large language models, incurring substantial computational costs.

By Teng Chen, Sheng Xu, Feixiang Guo, Xiaoyu Wang, Qingqing Gu, Hongyan Li, Luo Ji
arXiv AI
Sep 16

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

The paper introduces ORDER, a task‑conditioned retrieval‑augmented generation framework that dynamically adapts both indexing and retrieval strategies to each incoming query. It first clusters questions to learn cluster‑specific chunking, metadata filtering, and reranking settings, then routes queries to the appropriate pre‑built index via nearest‑centroid assignment. Additionally, a supervised query router predicts relevant collections and a Uniform Multi‑source Sampler distributes the retrieval budget evenly across selected sources, yielding superior performance on heterogeneous historical archives compared to existing RAG systems.

By Aur\'elien Pellet (LRE), Julien Perez, Marie Puren
arXiv AI
Sep 21

Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

The paper introduces DSRec, a dual‑interest sequential recommendation model that separates item representations into long‑term and short‑term semantic contexts. Long‑term embeddings capture stable preferences through historical aggregation, while short‑term embeddings focus on local session intent modulated by inter‑click time intervals. Each branch is processed by a distinct State Space Model— a full‑sequence Mamba for long‑term modeling and a time‑modulated SSM for short‑term dynamics— and a residual cross‑fusion mechanism aligns the two granularities while preserving their independence. Experiments on public benchmarks show that DSRec outperforms state‑of‑the‑art methods.

By Shuiying Liao, P. Y. Mok
arXiv AI
Jun 10

STORM: Stepwise Token Optimization with Reward-Guided Beam Search

arXiv:2606. 10621v1 Announce Type: cross Abstract: Modern retrieval increasingly relies on dense and learned-sparse neural models that are effective but require encoding the entire corpus into a specialized index, rebuilt whenever the model changes.

By Arthur Satouf, Giulio D'Erasmo, Yuxuan Zong, Habiboulaye Amadou Boubacar, Pablo Piantanida, Benjamin Piwowarski