Towards Data Science

What Are the Possibilities to Build Date Tables in Self-Service Environments?

For years, I created date tables with DAX code whenever I didn’t have a way to create them upstream of the data flow. Now I've realised there's another way to do it.

arXiv Machine Learning
Aug 12

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

arXiv:2607. 15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because temporal structure is easily lost under conventional tabular metrics.

By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv Machine Learning
Jul 20

Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

arXiv:2607. 15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and while sending entities along paths that no real entity followed.

By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
arXiv AI
Sep 1

Reviving our data foundations is the most disruptive step to data maturity

The article argues that for small‑to‑medium enterprises, the most disruptive yet essential step toward data maturity is to rebuild or strengthen a solid knowledge foundation layer. It stresses that this initiative must be evidence‑backed and minimally disruptive to current processes, and it proposes a low‑impact data strategy that adapts to evolving data flows. The authors emphasize that knowledge graph techniques will become indispensable in AI‑powered enterprises if designed modularly, dynamically, and cross‑functionally.

By Valentina Carapella, Ernesto Jimenez-Ruiz