arXiv:2606. 01840v1 Announce Type: new Abstract: Many difficult search problems cannot be solved by algorithms such as A* using only RAM.
By Yuki Suzuki, Alex Fukunaga
How Pandas chunking, Dask, and Polars help process millions of records when adding more compute isn't an option. The post What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
By Jiayan Yin
arXiv:2606. 08813v1 Announce Type: cross Abstract: We present HNTL (Hierarchical No-pointer Tangent-Local), the core vector indexing and candidate generation framework of the Aperon vector memory system.
By Yong Fu
arXiv:2609.20276v2 Announce Type: replace-cross
Abstract: Memoizing an expensive function of a sorted score vector is a data-structure problem before it is a numerical one: at a billion gridpoints, a...
By Tamal Maharaj
arXiv:2608.22141v1 Announce Type: new
Abstract: Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- faithful evid...
By Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao
arXiv:2608. 14648v1 Announce Type: cross Abstract: In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clustering: dimensionality reduction, quantization, and dimension pruning.
By Leonardo Kuffo, Peter Boncz