arXiv:2606. 24407v1 Announce Type: cross Abstract: We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity.
By Lorenzo Balzotti, Donatella Firmani, Luca Gagliardelli, Giovanni Simonini
We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interrogate such an oracle to resolve entities in a dataset whose size is far larger than a single batch, and where no batch is guaranteed to contain all records of any given entity.
The paper introduces a scalable product‑linking system that uses a retrieve‑then‑match cascade. First, a lightweight text cross‑encoder auto‑resolves the majority of merchant‑catalog product pairs with high precision, while an agentic multimodal vision‑language model handles the remaining ambiguous cases by inspecting images and performing web searches. This approach balances computational cost and accuracy, improving overall link coverage from 68% to 77% without requiring fine‑tuning of the agent.
By Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu, Charu Sareen, Kyle MacDonald
arXiv:2607. 09236v1 Announce Type: new Abstract: Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities, critical for privacy and safety.
By Amit Peleg, Naman Deep Singh, Naama Pearl, Bibhabasu Mohapatra, Matthias Hein
The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.
By Phanindra Reddy Madduru
OpenSanctions Pairs is the first large‑scale public benchmark for entity matching on sanctions and OSINT data, comprising 755,540 expert‑labeled pairs drawn from over 1 million entities across 293 source datasets and 45 jurisdictions. The dataset spans multiple languages and writing systems, inconsistent structures, and time‑varying provenance, making it far more heterogeneous than prior benchmarks. Baseline experiments show a rule‑based matcher achieving 91.3 % F1, GPT‑4o reaching 99.0 % F1, and a locally deployable open‑source model scoring 98.2 % F1, with complementary failure modes that highlight the need to focus on downstream pipeline components.
By Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt