arXiv AI By Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

Read the original on arXiv AI →

The paper introduces Hierarchical Hash Retrieval (HHR), a coarse‑to‑fine framework designed to improve hash‑based retrieval for large language models. HHR combines Geometry‑Aware Key Routing (GKR) to redistribute feature magnitudes and prune low‑logit keys, with Learned Hash Projection (LHP) to align Hamming distance with true query‑key relevance for fine‑grained retrieval. Experiments on diverse LLMs and benchmarks show that HHR outperforms existing methods, boosting LongBench scores by 1.10 points and achieving up to 3.30× decoding speedup at 128K context length for Llama‑3.1‑8B‑Instruct.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 6

Training-Free Hashing-Based Attention via Binary Principal Components

arXiv:2608. 04405v1 Announce Type: cross Abstract: Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches.

By Daohai Yu, Zhanpeng Zeng, Keyu Chen, Wenhao Li, Zhifeng Shen, Luxi Lin, Ruizhi Qiao, Xing Sun, Rongrong Ji
arXiv AI
Sep 10

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Matryoshka Hash Representations (MHR) propose a two‑stage quantization approach for retrieval‑augmented generation. First, a long binary code is learned; then, frozen, additional zero‑initialized residual adaptors are trained to produce searchable prefixes of varying byte budgets. Evaluated on MS MARCO and transferred to seven BEIR datasets, MHR achieves higher NDCG@10 and Recall@100 at 32‑byte budgets than baselines, especially in low‑budget regimes, and can also improve candidate shortlisting and graph‑index pruning.

By Peichun Hua, Yunming Xiao
Hugging Face Trending Papers
Aug 5

Training-Free Hashing-Based Attention via Binary Principal Components

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing.

arXiv AI
Jun 9

Projection and Quantisation: A Unifying View of Learning to Hash, from Random Projections to the RAG Era

arXiv:2510. 04127v2 Announce Type: replace-cross Abstract: Approximate nearest neighbour (ANN) search underpins large-scale retrieval, increasingly within the retrieval-augmented generation pipelines that ground large language models, yet the methods that address it have multiplied across communities until they are seldom read as a single field.

By Sean Moran