arXiv AI By Leah Woldemariam, Sudhanshu Garg, Taha Belkhouja, Charles Kim-Yip, Ali Sahami

Exploring Bottom-Up Clustering for Creating Semantic IDs

Read the original on arXiv AI →

The paper proposes a new algorithm for generating Semantic IDs that are both unique and preserve the structure of the original embedding space. By employing bottom‑up clustering, the method maintains local structure, leading to higher clustering quality. This improved structure enhances the utility of the Semantic IDs for downstream generative retrieval tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 26

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

arXiv:2512. 10388v3 Announce Type: replace-cross Abstract: Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture collaborative signals from historical user-item interactions.

By Ziwei Liu, Yejing Wang, Wanyu Wang, Wang Zejian, Qidong Liu, Zijian Zhang, Chong Chen, Wei Huang, Xiangyu Zhao
arXiv AI
2d ago

Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)

The paper introduces Graph-Informed Semantic IDs (GrIS), a framework that reframes Semantic ID construction as a recursive clustering problem on a graph combining semantic content and collaborative signals. GrIS generalises previous methods like RQ-VAE and RQ-KMeans by allowing explicit graph construction and hierarchical partitioning, and presents two implementations: RecDMoN and RQ-GAE. Experiments on real-world datasets show that GrIS outperforms collaborative-filtering aware state‑of‑the‑art models, achieving up to a 52% increase in Hit@10.

By Aleksei Medvedev, Alejandro Ariza-Casabona, Steven Derby, Gonzalo Fiz Pontiveros, Xinyang Shao, Florian Spiess
arXiv AI
4d ago

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

FineSID introduces a new quantization framework for semantic identifier learning in generative recommendation systems. By replacing the traditional Top‑1 hard assignment with a soft, differentiable approach, it distributes gradient updates across all codewords, leading to balanced codebook optimization and reduced identifier collisions. Experiments on public benchmarks show that FineSID improves codebook utilization and recommendation accuracy without relying on complex initialization strategies.

By Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang