arXiv:2602. 02025v2 Announce Type: replace-cross Abstract: ML models critically depend on feature quality, yet in real-world settings, useful features are often distributed across multiple relational tables rather than a single dataset.
By Serafeim Papadias, Kostas Patroumpas, Dimitrios Skoutas
arXiv:2606. 29532v1 Announce Type: cross Abstract: Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis.
By Christopher Gou, Aditya Banerjee, Jiaxuan Wang, Chunwei Liu
arXiv:2609.26658v1 Announce Type: cross
Abstract: Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit att...
By Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero, Sergi Nadal
arXiv:2610.01064v1 Announce Type: cross
Abstract: Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, bu...
By Sandipan De, Abhijit Chakraborty, Sambaran Bandyopadhyay, Vivek Gupta
arXiv:2601.13111v3 Announce Type: replace-cross
Abstract: Realistic text-to-SQL workflows often require joining multiple tables. As a result, accurately retrieving the relevant set of tables becomes...
By Hassan Soliman, Vivek Gupta, Dan Roth, Iryna Gurevych
arXiv:2607. 03659v1 Announce Type: cross Abstract: Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been evaluated under inconsistent experimental protocols, making fair comparison difficult.
By Kazi F. Akhter, Bharath Ajendla, Manar D. Samad
arXiv:2506. 18421v3 Announce Type: replace-cross Abstract: The majority of data in businesses and industries is stored in tables, databases, and data warehouses.
By Ce Li, Xiaofan Liu, Zhiyan Song, Ce Chi, Boshen Shi, Chen Zhao, Guanguang Chang, Zhendong Wang, Kexin Yang, Xing Wang, Chao Deng, Junlan Feng
The paper introduces FlockMTL, an extension for database management systems that deeply integrates large language models and retrieval‑augmented generation into DuckDB. It provides model‑driven scalar and aggregate functions, cost‑based optimizations like batching and caching, and new SQL DDL abstractions (PROMPT and MODEL) to treat LLMs as first‑class schema objects. These features aim to simplify the development of knowledge‑intensive analytical applications by reducing the effort required to orchestrate heterogeneous data systems and manage LLM context.
By Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi
arXiv:2602. 10441v2 Announce Type: replace Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data.
By Feiyu Pan, Tianbin Zhang, Aoqian Zhang, Yu Sun, Zheng Wang, Lixing Chen, Li Pan, Jianhua Li
The paper introduces ColRel, a two-stage approach that uses large language models to discover relationships between columns in data lakes. It first creates column embeddings from available metadata and data at ingestion, then refines these embeddings with business dictionaries to generate concise natural-language descriptions. Experiments on public benchmarks and an industrial ERP dataset demonstrate ColRel’s effectiveness, especially in scenarios with weak signals and semantically related columns.
By Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2)
arXiv:2609.20886v1 Announce Type: cross
Abstract: Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tab...
By Chuxuan Hu, Yeye He, Penny Zhou, Wee Hyong Tok, Daniel Kang, Surajit Chaudhuri
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream workflows rely on normalized schemas, entity identities, keys, cross-table relationships, and integrity constraints for analytics, compliance, auditing, and SQL-backed decision making.