arXiv Machine Learning By Zeyu Zhang, Xue Li, Iacer Calixto, Paul Groth, Sebastian Schelter

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Read the original on arXiv Machine Learning →

arXiv:2607. 24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 24

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv:2607. 20465v1 Announce Type: new Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end.

By Hao Liang, Qifeng Cai, Yibo Lin, Jianzhuo Du, Qifeng Xia, Sizhe Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, Bohan Zeng, Ruichuan An, Conghui He, Wentao Zhang