arXiv Machine Learning

Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

arXiv:2607. 15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and while sending entities along paths that no real entity followed.

arXiv Machine Learning
Aug 12

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

arXiv:2607. 15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because temporal structure is easily lost under conventional tabular metrics.

By Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv Machine Learning
Jun 9

Disjoint Generation of Synthetic Data

arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.

By Anton Danholt Lautrup, Muhammad Rajabinasab, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp
arXiv AI
Aug 11

Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations

arXiv:2608. 08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production traffic.

By Michael Levit, Josh Ledgard, Haoyu Dong, Vishwas Suryanarayanan, Eyal Kolman, Sharon Tan, Qiang Gan, Vishal Chowdhary