arXiv AI

A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests

The paper introduces SCR-MF, a two‑stage workflow for single‑cell RNA sequencing imputation that first detects dropout events with scRecover and then imputes missing values using the non‑parametric missForest algorithm. Benchmarking on public and simulated datasets shows that SCR‑MF delivers robust, interpretable results that match or surpass existing methods while maintaining biological fidelity. Runtime analysis indicates that SCR‑MF balances accuracy with computational efficiency, making it well suited for mid‑scale single‑cell studies.

arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv AI
Sep 15

Towards a knowledge-enhanced single-cell foundation model

The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.

By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca
arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv Machine Learning
Jun 9

Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.

By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo