Towards Data Science

Avoiding Entity Key Drift in a Data Lake: Step 2, When Fuzzy Matching Stops Working

The article discusses a matcher designed to clean up residual issues after normalization in a data lake. Testing revealed that no version of the matcher could be made fully safe, leading to its abandonment. The post then outlines the architecture that remained after the matcher was set aside.

arXiv Machine Learning
Aug 27

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

OpenSanctions Pairs is the first large‑scale public benchmark for entity matching on sanctions and OSINT data, comprising 755,540 expert‑labeled pairs drawn from over 1 million entities across 293 source datasets and 45 jurisdictions. The dataset spans multiple languages and writing systems, inconsistent structures, and time‑varying provenance, making it far more heterogeneous than prior benchmarks. Baseline experiments show a rule‑based matcher achieving 91.3 % F1, GPT‑4o reaching 99.0 % F1, and a locally deployable open‑source model scoring 98.2 % F1, with complementary failure modes that highlight the need to focus on downstream pipeline components.

By Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt
Google AI Blog
Mar 11, 2024

Chain-of-table: Evolving tables in the reasoning chain for table understanding

Posted by Zilong Wang, Student Researcher, and Chen-Yu Lee, Research Scientist, Cloud AI Team People use tables every day to organize and interpret complex information in a structured, easily accessible format. Due to the ubiquity of such tables, reasoning over tabular data has long been a central topic in natural language processing (NLP).

By Google AI