arXiv:2607. 12188v1 Announce Type: new Abstract: Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants.
By Navnit Shukla
Matryoshka Hash Representations (MHR) propose a two‑stage quantization approach for retrieval‑augmented generation. First, a long binary code is learned; then, frozen, additional zero‑initialized residual adaptors are trained to produce searchable prefixes of varying byte budgets. Evaluated on MS MARCO and transferred to seven BEIR datasets, MHR achieves higher NDCG@10 and Recall@100 at 32‑byte budgets than baselines, especially in low‑budget regimes, and can also improve candidate shortlisting and graph‑index pruning.
By Peichun Hua, Yunming Xiao
Spruce is a system that enables scalable private outsourced retrieval by learning compact binary embeddings and using efficient Hamming-distance computation under a two‑server multi‑party computation protocol. It replaces costly corpus‑wide embedding scoring with a fixed‑radius protocol that avoids multi‑round candidate selection, and introduces private cluster pruning and a one‑core dealer to reduce computation and eliminate OT preprocessing bottlenecks. Across corpora of 383K–5.42M documents, Spruce maintains original search quality while achieving up to 31.5× higher throughput and reducing query times to a few seconds.
By Peichun Hua, Yunming Xiao
arXiv:2608. 05127v1 Announce Type: cross Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging.
By Adel Javanmard, David P. Woodruff, Vahab Mirrokni
MOMAT (Mixture of Multiple Atlases) is a hardware‑enhanced safety framework designed to defend quantized large language models (qLLMs) on edge devices against jailbreak attacks. It uses a collection of semantic atlases—each containing harmful or benign sample clusters and policy templates—to perform domain‑localized Retrieval‑Augmented Generation. A lightweight Mixture of Experts detector evaluates top‑k similarity features retrieved by a Compute‑in‑Memory (CiM) accelerated engine, achieving a 4.69 × 10⁶‑fold speedup and a 2.5 × 10⁵‑fold energy reduction compared to DRAM‑based baselines while matching state‑of‑the‑art defense performance.
By Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin
Spruce is a system that enables secure, private retrieval of large document collections outsourced to untrusted clouds by learning compact binary embeddings that preserve search quality while drastically reducing computation and communication. It replaces expensive corpus-wide embedding scoring with efficient Hamming-distance calculations under a two-server multi-party computation protocol, and introduces a fixed-radius protocol, private cluster pruning, and a one-core dealer to further cut latency and bandwidth usage. Across corpora ranging from 383K to 5.42M documents, Spruce maintains original search quality, achieving up to 6.7× faster full scans and 22.9× speedups with pruning, while retaining over 94% of the original NDCG.