arXiv Machine Learning

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

arXiv:2607. 26298v1 Announce Type: new Abstract: We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature.

Hugging Face Trending Papers
Jun 23

Entity Resolution via Batched Oracle Queries

We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interrogate such an oracle to resolve entities in a dataset whose size is far larger than a single batch, and where no batch is guaranteed to contain all records of any given entity.

arXiv Computation and Language
Aug 27

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

The paper introduces a scalable product‑linking system that uses a retrieve‑then‑match cascade. First, a lightweight text cross‑encoder auto‑resolves the majority of merchant‑catalog product pairs with high precision, while an agentic multimodal vision‑language model handles the remaining ambiguous cases by inspecting images and performing web searches. This approach balances computational cost and accuracy, improving overall link coverage from 68% to 77% without requiring fine‑tuning of the agent.

By Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu, Charu Sareen, Kyle MacDonald
arXiv AI
Sep 3

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.

By Phanindra Reddy Madduru
arXiv Machine Learning
Aug 27

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

OpenSanctions Pairs is the first large‑scale public benchmark for entity matching on sanctions and OSINT data, comprising 755,540 expert‑labeled pairs drawn from over 1 million entities across 293 source datasets and 45 jurisdictions. The dataset spans multiple languages and writing systems, inconsistent structures, and time‑varying provenance, making it far more heterogeneous than prior benchmarks. Baseline experiments show a rule‑based matcher achieving 91.3 % F1, GPT‑4o reaching 99.0 % F1, and a locally deployable open‑source model scoring 98.2 % F1, with complementary failure modes that highlight the need to focus on downstream pipeline components.

By Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt
arXiv Machine Learning
Sep 7

Corporate-Family Resolution Is Not a String-Matching Problem: A Public Benchmark Stratified by Name Visibility

The paper introduces CorpFam, a public benchmark for corporate-family resolution comprising 54,864 candidate pairs across 10,307 families derived from 6.6 million US federal award records. It shows that traditional entity-matching methods perform poorly on pairs where name visibility is low, with the best matcher recovering only 4.2% of such invisible links and blocking schemes failing to identify most of them. The study argues that the real challenge lies in candidate generation rather than ranking, and validates the benchmark’s links against SEC subsidiary schedules, confirming their authenticity.

By Harshit Gupta
arXiv Computation and Language
Sep 23

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.

By Samuele Garda, Ulf Leser
arXiv AI
Jul 21

Accurate and Efficient Long-Term Memory for LLM Agents

arXiv:2607. 16211v1 Announce Type: new Abstract: LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment.

By Zicheng Zhao, Xinyang Guo, Luyao Lv, Menghan Wang, Ming Li, Shuaicheng Li