Flow Matching with Missing Data
arXiv:2607. 28698v1 Announce Type: new Abstract: Flow matching assumes fully observed training data, which many real-world applications rarely provide.
arXiv:2607. 23295v1 Announce Type: cross Abstract: In real-world machine learning applications, incomplete observations create a fundamental challenge.
arXiv:2607. 28698v1 Announce Type: new Abstract: Flow matching assumes fully observed training data, which many real-world applications rarely provide.
The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
arXiv:2607. 07640v1 Announce Type: cross Abstract: Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within the corrupted input sequence.
arXiv:2607. 03641v1 Announce Type: cross Abstract: The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold.
arXiv:2607. 06930v1 Announce Type: cross Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis.
arXiv:2609.37632v1 Announce Type: cross Abstract: Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent perform...
arXiv:2504. 15388v3 Announce Type: replace-cross Abstract: In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique.
arXiv:2606. 16484v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language.
arXiv:2606. 15743v1 Announce Type: new Abstract: This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature embeddings in a task-independent manner before supervised prediction.
arXiv:2607. 08915v1 Announce Type: new Abstract: Missing data is ubiquitous in real-world datasets.
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
arXiv:2602.01437v2 Announce Type: replace-cross Abstract: The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address...