arXiv Machine Learning By Yuyu Liu, Jiannan Yang, Ziyang Yu, Weishen Pan, Fei Wang, Tengfei Ma

Efficient Imputation for Patch-based Missing Single-cell Data via Cluster-regularized Optimal Transport

Read the original on arXiv Machine Learning →

arXiv:2601. 14653v3 Announce Type: replace Abstract: Missing data in single-cell sequencing datasets poses significant challenges for extracting meaningful biological insights.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 21

A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests

The paper introduces SCR-MF, a two‑stage workflow for single‑cell RNA sequencing imputation that first detects dropout events with scRecover and then imputes missing values using the non‑parametric missForest algorithm. Benchmarking on public and simulated datasets shows that SCR‑MF delivers robust, interpretable results that match or surpass existing methods while maintaining biological fidelity. Runtime analysis indicates that SCR‑MF balances accuracy with computational efficiency, making it well suited for mid‑scale single‑cell studies.

By Ali Anaissi, Deshao Liu, Yuanzhe Jia, Weidong Huang, Widad Alyassine, Junaid Akram
arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi