Towards Data Science By Thomas Reid

I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance.

Read the original on Towards Data Science →

The article reports on a benchmark that reduced 1,000 Apache Iceberg files to just six, then measured how this consolidation affected query performance across three SQL workloads. It details the methodology and results of the experiment, highlighting changes in execution speed and resource usage. The findings illustrate the trade‑offs between file count and query efficiency in large data systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

Towards Data Science
Aug 25

One Document Type, a Million Files: Structured Extraction into the SQL Table RAG Queries

The article discusses a case study in enterprise document intelligence where a single document type contains a million files. It outlines a workflow that takes about an hour with two people to extract six to ten structured fields, emphasizing the importance of identifying two key signals that distinguish a valid column from one that could break a filter later. The focus is on converting unstructured documents into a structured SQL table for Retrieval-Augmented Generation (RAG) queries.

By Angela and Kezhan Shi
arXiv AI
Aug 18

Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps

arXiv:2608. 16309v1 Announce Type: cross Abstract: Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations and dynamic pruning mechanisms.

By Zirui Song, Yuye Zhu, Yang Yang
Simon Willison
Sep 7

Creepy crawlies

Konstantin Ryabitsev highlights the growing problem of abusive web crawlers that consume excessive CPU resources on git.kernel.org, the official Git repository for the Linux kernel. He notes that at any given moment, 14 CPU cores across five geo‑distributed nodes are dedicated solely to rendering git commits as HTML for these scrapers, surpassing the CPU usage for all legitimate access such as git clones. This issue raises concerns for services like Datasette, which also serve large numbers of crawlable web pages.