arXiv:2605.07938v2 Announce Type: replace
Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular funct...
By Sachini Weerasekara, Natasha Darras, Sagar Kamarthi, Colles Price, Jacqueline Isaacs
arXiv:2511. 02986v2 Announce Type: replace-cross Abstract: Computational modeling of single-cell gene expression is crucial for understanding cellular processes, but generating realistic expression profiles remains a major challenge.
By Giovanni Palla, Sudarshan Babu, Payam Dibaeinia, James D. Pearce, Donghui Li, Aly A. Khan, Theofanis Karaletsos, Jakub M. Tomczak
arXiv:2602. 15253v2 Announce Type: replace Abstract: Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transformers, yet their existence in single-cell genomics remains largely unexplored.
By Ihor Kendiukhov
The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.
By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca
arXiv:2608. 00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values.
By Jiaqi Xiong, Yuntao hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
PopPert is a framework that models population-level joint gene expression distributions to predict transcriptional responses to perturbations in single-cell RNA sequencing data. By using a low‑rank Gaussian Copula, it captures gene co‑expression patterns and eliminates the need for cell‑to‑cell correspondence, thereby reducing sensitivity to single‑cell noise. Across multiple benchmarks, PopPert outperforms existing methods in differential expression recovery, perturbation effect estimation, and distribution matching, demonstrating the effectiveness of population‑level joint distribution learning for unpaired single‑cell data.
By Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai
arXiv:2606. 27752v1 Announce Type: new Abstract: Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions.
By Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv:2607. 19426v1 Announce Type: cross Abstract: Single-cell datasets are increasingly costly to store, audit, and reuse for model training.
By Yaodi Luo, Peize He, Bowen Han, Lingbei Mengg
arXiv:2607. 19426v2 Announce Type: replace-cross Abstract: Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training.
By Yaodi Luo, Peize He, Lingbei Meng, Bowen Han, Zheng Lu, Jianqing Zhu, Lian Zhang
CellRFT is a reinforcement fine‑tuning framework designed to improve single‑cell perturbation modeling by directly optimizing biological evaluation metrics. It employs policy‑gradient methods to learn from non‑differentiable biological rewards and aggregates multiple rewards hierarchically. Experiments show that CellRFT enhances perturbation prediction across various pretrained models and reveals interactions between different biological criteria, suggesting new ways to shape model behavior and evaluation design.
By Jie Yan, Li Liu, Hanze Guo, Jiaxin Hu, Houxin He, Xiaoning Qi, Haoran Wang, Cong Li, Zhong-Yuan Zhang, Yong Wang
arXiv:2606. 13713v1 Announce Type: cross Abstract: Predicting cellular transcriptional responses to genetic perturbations is a central problem in single-cell biology, especially in the zero-shot setting where the perturbed gene or gene combination is unseen during training.
By Wei Zhang, Xun Jiang, Yuesi Xi, Ming Tang