arXiv AI By Navnit Shukla, Kamal Pandey, Omsankar Tiwari

TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization

Read the original on arXiv AI →

arXiv:2607. 16973v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two underexplored challenges: (1) trained codebook quantizers may expose corpus statistics during index construction, creating a leakage channel in multi-tenant deployments, and (2) post-hoc filtering for tenant isolation degrades recall on selective queries.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 15

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

arXiv:2607. 12188v1 Announce Type: new Abstract: Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants.

By Navnit Shukla
arXiv AI
Sep 10

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Matryoshka Hash Representations (MHR) propose a two‑stage quantization approach for retrieval‑augmented generation. First, a long binary code is learned; then, frozen, additional zero‑initialized residual adaptors are trained to produce searchable prefixes of varying byte budgets. Evaluated on MS MARCO and transferred to seven BEIR datasets, MHR achieves higher NDCG@10 and Recall@100 at 32‑byte budgets than baselines, especially in low‑budget regimes, and can also improve candidate shortlisting and graph‑index pruning.

By Peichun Hua, Yunming Xiao
arXiv Machine Learning
Sep 4

Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings

Spruce is a system that enables scalable private outsourced retrieval by learning compact binary embeddings and using efficient Hamming-distance computation under a two‑server multi‑party computation protocol. It replaces costly corpus‑wide embedding scoring with a fixed‑radius protocol that avoids multi‑round candidate selection, and introduces private cluster pruning and a one‑core dealer to reduce computation and eliminate OT preprocessing bottlenecks. Across corpora of 383K–5.42M documents, Spruce maintains original search quality while achieving up to 31.5× higher throughput and reducing query times to a few seconds.

By Peichun Hua, Yunming Xiao
arXiv AI
2d ago

MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs

MOMAT (Mixture of Multiple Atlases) is a hardware‑enhanced safety framework designed to defend quantized large language models (qLLMs) on edge devices against jailbreak attacks. It uses a collection of semantic atlases—each containing harmful or benign sample clusters and policy templates—to perform domain‑localized Retrieval‑Augmented Generation. A lightweight Mixture of Experts detector evaluates top‑k similarity features retrieved by a Compute‑in‑Memory (CiM) accelerated engine, achieving a 4.69 × 10⁶‑fold speedup and a 2.5 × 10⁵‑fold energy reduction compared to DRAM‑based baselines while matching state‑of‑the‑art defense performance.

By Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin
Hugging Face Trending Papers
Sep 3

Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings

Spruce is a system that enables secure, private retrieval of large document collections outsourced to untrusted clouds by learning compact binary embeddings that preserve search quality while drastically reducing computation and communication. It replaces expensive corpus-wide embedding scoring with efficient Hamming-distance calculations under a two-server multi-party computation protocol, and introduces a fixed-radius protocol, private cluster pruning, and a one-core dealer to further cut latency and bandwidth usage. Across corpora ranging from 383K to 5.42M documents, Spruce maintains original search quality, achieving up to 6.7× faster full scans and 22.9× speedups with pruning, while retaining over 94% of the original NDCG.