arXiv AI

Exploring Bottom-Up Clustering for Creating Semantic IDs

The paper proposes a new algorithm for generating Semantic IDs that are both unique and preserve the structure of the original embedding space. By employing bottom‑up clustering, the method maintains local structure, leading to higher clustering quality. This improved structure enhances the utility of the Semantic IDs for downstream generative retrieval tasks.

arXiv AI
Jun 26

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

arXiv:2512. 10388v3 Announce Type: replace-cross Abstract: Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture collaborative signals from historical user-item interactions.

By Ziwei Liu, Yejing Wang, Wanyu Wang, Wang Zejian, Qidong Liu, Zijian Zhang, Chong Chen, Wei Huang, Xiangyu Zhao
arXiv AI
2d ago

Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)

The paper introduces Graph-Informed Semantic IDs (GrIS), a framework that reframes Semantic ID construction as a recursive clustering problem on a graph combining semantic content and collaborative signals. GrIS generalises previous methods like RQ-VAE and RQ-KMeans by allowing explicit graph construction and hierarchical partitioning, and presents two implementations: RecDMoN and RQ-GAE. Experiments on real-world datasets show that GrIS outperforms collaborative-filtering aware state‑of‑the‑art models, achieving up to a 52% increase in Hit@10.

By Aleksei Medvedev, Alejandro Ariza-Casabona, Steven Derby, Gonzalo Fiz Pontiveros, Xinyang Shao, Florian Spiess
arXiv AI
4d ago

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

FineSID introduces a new quantization framework for semantic identifier learning in generative recommendation systems. By replacing the traditional Top‑1 hard assignment with a soft, differentiable approach, it distributes gradient updates across all codewords, leading to balanced codebook optimization and reduced identifier collisions. Experiments on public benchmarks show that FineSID improves codebook utilization and recommendation accuracy without relying on complex initialization strategies.

By Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang
arXiv AI
Sep 1

Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval

The paper introduces CHAP, a personalized generative retrieval framework that aligns query semantics with item representations through a hierarchical semantic alignment module and models user behavior using both discrete Semantic IDs and continuous representations. It also proposes a Residual Cascading Generation mechanism to reduce inference latency by limiting the Transformer decoder to a single pass. Experiments on multiple datasets and online A/B tests show that CHAP outperforms existing methods, demonstrating its practical value.

By Gaoming Zhang, Angqing Jiang, Jianchun Song, Kena Qi, Dayao Chen, Wei Lin, Defu Lian