arXiv AI By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Denis Kleyko, Vaclav Snasel

Derandomizing Dense Binary Hypervector Codebooks for Quantized Scalars

Read the original on arXiv AI →

The paper introduces a transition-based derandomization framework for dense binary hypervector codebooks used in hyperdimensional computing. It targets two similarity families—exponential and linear decay with scalar separation—and separates the similarity law, derandomization variant, and generator construction. The authors formalize variants that constrain initial Hamming weight, update-count variability, and update balance, deriving exact finite-dimensional expressions for bias, variance, and RMS error, and validate the theory with simulations to guide practical codebook design.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

FineSID introduces a new quantization framework for semantic identifier learning in generative recommendation systems. By replacing the traditional Top‑1 hard assignment with a soft, differentiable approach, it distributes gradient updates across all codewords, leading to balanced codebook optimization and reduced identifier collisions. Experiments on public benchmarks show that FineSID improves codebook utilization and recommendation accuracy without relying on complex initialization strategies.

By Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang
arXiv Computation and Language
Sep 18

D-Quant: Driftable Entropy Coding for KV Cache Quantization

The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.

By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu