The article discusses a case study in enterprise document intelligence where a single document type contains a million files. It outlines a workflow that takes about an hour with two people to extract six to ten structured fields, emphasizing the importance of identifying two key signals that distinguish a valid column from one that could break a filter later. The focus is on converting unstructured documents into a structured SQL table for Retrieval-Augmented Generation (RAG) queries.
By Angela and Kezhan Shi
arXiv:2608. 08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance.
By Jannic Cutura, Subash Prakash
arXiv:2608. 16309v1 Announce Type: cross Abstract: Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations and dynamic pruning mechanisms.
By Zirui Song, Yuye Zhu, Yang Yang
arXiv:2607. 11942v1 Announce Type: cross Abstract: KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol.
By Daming Luo, Christy Liang, Junyu Xuan
Increasing context size in RAG systems doesn’t improve accuracy for aggregation tasks—it makes errors harder to detect. In this article, I benchmark retrieval-based pipelines against a deterministic full-scan engine across 100,000 rows and show why computation queries must be routed away from RAG entirely.
By Emmimal P Alexander
Konstantin Ryabitsev highlights the growing problem of abusive web crawlers that consume excessive CPU resources on git.kernel.org, the official Git repository for the Linux kernel. He notes that at any given moment, 14 CPU cores across five geo‑distributed nodes are dedicated solely to rendering git commits as HTML for these scrapers, surpassing the CPU usage for all legitimate access such as git clones. This issue raises concerns for services like Datasette, which also serve large numbers of crawlable web pages.
A practical reproduction of three retrieval baselines, including the crashes, fixes, and score checks that matter for RAG systems. The post How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook appeared first on Towards Data Science .
By Abdullahi Dattijo
arXiv:2609.37911v1 Announce Type: cross
Abstract: Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier...
By Ryan C. Barron, Cade W. Trotter, Maksim E. Eren, Kim {\O}. Rasmussen, Liz D. Miller, Benjamin J. Migliori
arXiv:2607. 28880v1 Announce Type: cross Abstract: Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3.
By Debopam Sanyal, Hongjie Chen, Alexey Tumanov, Joshua Kimball
arXiv:2609.01525v1 Announce Type: cross
Abstract: A durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data....
By Gene Zhang
arXiv:2609.35918v1 Announce Type: cross
Abstract: Finding the right files is an early challenge for coding agents. We test whether a language model can follow directory and file names to find annotat...
By Manoj Bajaj
The true bottleneck was never the analysis. The post BI Is Dead, Long Live BI appeared first on Towards Data Science .
By Mahdi Karabiben