arXiv AI By Johnny Greco, Nabin Mulepati, Andre Manoel, Eric Tramel, Kirit Thadaka, Mike Knepper, Dhruv Nathawani, Dane Corneil, Yev Meyer, Alex Watson, Maarten Van Segbroeck

NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

Read the original on arXiv AI →

NeMo Data Designer (NDD) is an open‑source framework for generating multimodal synthetic data. It uses a declarative configuration format that lets users define dataset columns—text, code, structured outputs, images, embeddings, and statistical samplers—to steer diversity. The system supports a preview‑and‑revision workflow, dependency resolution, and retry logic, and can be extended via plugins. Case studies demonstrate its use for structured, agentic, multimodal, and domain‑specialized tasks, including datasets for Nemotron model development and enterprise deployments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 7

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

arXiv:2605. 17758v2 Announce Type: replace Abstract: Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information.

By Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran, Tu Nguyen, Arshia Harish Puthran, Muhjaazee Love, Aadi Sharma, Mahdi Bagheri, Ian Harris, Amir M. Rahmani
arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv AI
Jul 21

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

arXiv:2607. 16617v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts.

By Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang