arXiv:2607. 15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because temporal structure is easily lost under conventional tabular metrics.
By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
arXiv:2512.00329v2 Announce Type: replace-cross
Abstract: Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task...
By Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta, Vivek Gupta
One of the most important concepts in DAX is lineage. It’s about the information on where something comes from.
By Salvatore Cagliari
arXiv:2606. 02433v1 Announce Type: cross Abstract: The rapid development of LLMs has significantly advanced tabular question answering, but most systems cannot perform future-oriented numerical prediction.
By Zhensheng Wang, Xiaole Liu, Wenmian Yang, Kun Zhou, Yiquan Zhang, Weijia Jia
arXiv:2606. 30452v1 Announce Type: new Abstract: Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks.
By Myung Jun Kim, Maximilian Schambach, Frank Essenberger, Andre Sres, Johannes H\"ohne
Release: alchemy-utils 0. 1a0 I've long pondered what a database agnostic version of my sqlite-utils Python library and CLI utility might look like.
Starting with a local Parquet file, then joining it to data stored in the cloud
The post Building a Data Lakehouse with DuckDB and DuckLake appeared first on Towards Data Science.
By Thomas Reid
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv:2607. 15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and while sending entities along paths that no real entity followed.
By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
Tabular foundation models predict the missing column of any spreadsheet zero-shot, the way an LLM completes text — and on the TabArena benchmark they now sit above fully tuned gradient-boosted trees. An introduction to how they work, an independent reproduction of the strongest open one, and a map of where XGBoost still wins.
By Sean Moran
The article argues that for small‑to‑medium enterprises, the most disruptive yet essential step toward data maturity is to rebuild or strengthen a solid knowledge foundation layer. It stresses that this initiative must be evidence‑backed and minimally disruptive to current processes, and it proposes a low‑impact data strategy that adapts to evolving data flows. The authors emphasize that knowledge graph techniques will become indispensable in AI‑powered enterprises if designed modularly, dynamically, and cross‑functionally.
By Valentina Carapella, Ernesto Jimenez-Ruiz