GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv:2607. 29129v1 Announce Type: new Abstract: Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline.
By Mohammad Sadeq Abolhasani, Viswanath Ganapathy
arXiv:2602. 13697v2 Announce Type: replace-cross Abstract: Relational databases (RDBs) contain vast amounts of heterogeneous tabular information that can be exploited for predictive modeling purposes.
By Linjie Xu, Yanlin Zhang, Quan Gan, Minjie Wang, David Wipf
The paper introduces FlockMTL, an extension for database management systems that deeply integrates large language models and retrieval‑augmented generation into DuckDB. It provides model‑driven scalar and aggregate functions, cost‑based optimizations like batching and caching, and new SQL DDL abstractions (PROMPT and MODEL) to treat LLMs as first‑class schema objects. These features aim to simplify the development of knowledge‑intensive analytical applications by reducing the effort required to orchestrate heterogeneous data systems and manage LLM context.
By Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi
arXiv:2608. 16319v1 Announce Type: new Abstract: This first release of Prior Labs in relational learning shows our continued commitment to open science.
By Adrian Hayler, Klemens Fl\"oge, Alan Arazi, Rishabh Ranjan, Jure Leskovec, Felix Birkel, Brendan Roof, Anurag Garg, Kristina Collins, Lydia Sidhoum, Jonas K\"ubler, Siyuan Guo, Oscar Key, Jan Hendrik Metzen, Rylee Grace, David Salinas, Arthur Cahu, Simon Bing, Benjamin J\"ager, Tuana \c{C}elik, Mihir Manium, Vitor Monteiro, Jake Robertson, Jerry Chen, Eliott Kalfon, Tom\'as Pereda, Lilly Wehrhahn, Dominik Safaric, Tobias Schroeder, Georg Grab, Diana Kriuchkova, Clara Cornu, Philipp Singer, Nick Erickson, Vahid Balazadeh, Marie Salmon, Simone Alessi, K\"ur\c{s}at Kaya, Philipp Jund, L\'eo Grinsztajn, Yann LeCun, Bernhard Sch\"olkopf, Madelon Hulsebos, Lennart Purucker, Sauraj Gambhir, Frank Hutter, Noah Hollmann
RelICL: Training-free Relational Learning with Tabular Foundation Models proposes a new method for relational learning that addresses two key issues of deep feature synthesis—feature explosion and interaction blindness—by propagating and fusing information step by step through the schema graph using a tabular foundation model. The approach retains the benefits of DFS while improving scalability and performance. Experiments on RelBench tasks show that RelICL performs on par with the strongest DFS-based approach.
By Simon Forbat, Rainer Gemulla