arXiv AI By Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2)

Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case

Read the original on arXiv AI →

The paper introduces ColRel, a two-stage approach that uses large language models to discover relationships between columns in data lakes. It first creates column embeddings from available metadata and data at ingestion, then refines these embeddings with business dictionaries to generate concise natural-language descriptions. Experiments on public benchmarks and an industrial ERP dataset demonstrate ColRel’s effectiveness, especially in scenarios with weak signals and semantically related columns.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP

SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.

By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang