The paper introduces ColRel, a two-stage approach that uses large language models to discover relationships between columns in data lakes. It first creates column embeddings from available metadata and data at ingestion, then refines these embeddings with business dictionaries to generate concise natural-language descriptions. Experiments on public benchmarks and an industrial ERP dataset demonstrate ColRel’s effectiveness, especially in scenarios with weak signals and semantically related columns.
By Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2)
arXiv:2610.00817v1 Announce Type: cross
Abstract: Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstrea...
By Sandipan De, Jin Wang, Vivek Gupta
SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.
By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv:2606. 28601v1 Announce Type: cross Abstract: Recent progress in Text-to-SQL has been driven by stronger language models and prompting strategies, yet performance on real enterprise benchmarks such as Spider 2.
By Jingwen Liu, Weibin Liao, Xin Gao, Junfeng Zhao, Yasha Wang
arXiv:2609.26658v1 Announce Type: cross
Abstract: Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit att...
By Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero, Sergi Nadal
Retrieval in the SQL setting has largely been studied as the task of finding, within a large collection of SQL statements, the statement that answers a natural-language question. At scale, however, a more fundamental retrieval problem precedes generation: schema retrieval, identifying the tables and columns a question requires in a database that may contain thousands of them, far more than fit in a model's context.
arXiv:2602. 16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments.
By Bowen Cao, Weibin Liao, Yushi Sun, Dong Fang, Haitao Li, Wai Lam
arXiv:2606. 07843v1 Announce Type: cross Abstract: Schema matching, a critical task for integrating data from diverse sources, seeks to identify correspondences between columns across different schemas.
By Leonard Traeger, Enas Khwaileh, Andreas Behrend, George Karabatis
arXiv:2608. 14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links.
By Yiming Zhang, Koji Tsuda
arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.
By Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
By So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen
arXiv:2607. 00008v1 Announce Type: cross Abstract: Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and complex.
By Sin Yu Bonnie Ho, Arlie Coles, Erik Larsson, Eric Marshall, Nathan Bodenstab, Paul Vozila