arXiv Machine Learning By Zeyu Zhang, Xue Li, Iacer Calixto, Paul Groth, Sebastian Schelter

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Read the original on arXiv Machine Learning →

arXiv:2607. 24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

OpenSanctions Pairs is the first large‑scale public benchmark for entity matching on sanctions and OSINT data, comprising 755,540 expert‑labeled pairs drawn from over 1 million entities across 293 source datasets and 45 jurisdictions. The dataset spans multiple languages and writing systems, inconsistent structures, and time‑varying provenance, making it far more heterogeneous than prior benchmarks. Baseline experiments show a rule‑based matcher achieving 91.3 % F1, GPT‑4o reaching 99.0 % F1, and a locally deployable open‑source model scoring 98.2 % F1, with complementary failure modes that highlight the need to focus on downstream pipeline components.

By Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt
arXiv Machine Learning
Sep 2

Can LLMs Use Relational Transformer Embeddings?

The paper investigates whether large language models (LLMs) can leverage frozen relational‑transformer embeddings by injecting them as soft tokens. Using a learned MLP projection and LoRA adaptation, the authors fine‑tune Qwen3.5‑4B on chain‑of‑thought reasoning traces and group‑based reinforcement learning, then evaluate on ten binary classification tasks across six RelBench databases. The hybrid approach consistently underperforms the standalone relational transformer, showing sensitivity to serialization format, token budget, and RL stability, leading the authors to conclude that stronger alignment objectives and schema‑aware design are needed for reliable relational prediction.

By Francisco Galuppo Azevedo, Clarissa Lima Loures