arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.
By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.
By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca
arXiv:2606. 14734v1 Announce Type: cross Abstract: Motivation: Gene regulatory network inference from single-cell RNA sequencing (scRNA-seq) data is important for uncovering cell-state-specific transcriptional programs.
By Ziyang Dong, Shanwen Tan, Hengchuang Yin, Wei Liu, Yifan Wang, Siyu Yi, Jiancheng Lv, Wei Ju
arXiv:2608. 00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values.
By Jiaqi Xiong, Yuntao hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
arXiv:2512. 17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling.
By Daphn\'e Chopard, Jorge da Silva Gon\c{c}alves, Irene Cannistraci, Thomas M. Sutter, Julia E. Vogt
arXiv:2605.07938v2 Announce Type: replace
Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular funct...
By Sachini Weerasekara, Natasha Darras, Sagar Kamarthi, Colles Price, Jacqueline Isaacs
PopPert is a framework that models population-level joint gene expression distributions to predict transcriptional responses to perturbations in single-cell RNA sequencing data. By using a low‑rank Gaussian Copula, it captures gene co‑expression patterns and eliminates the need for cell‑to‑cell correspondence, thereby reducing sensitivity to single‑cell noise. Across multiple benchmarks, PopPert outperforms existing methods in differential expression recovery, perturbation effect estimation, and distribution matching, demonstrating the effectiveness of population‑level joint distribution learning for unpaired single‑cell data.
By Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai
The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.
By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park
arXiv:2607. 16053v1 Announce Type: cross Abstract: Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints.
By Claudia Skok Gibbs
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence.
arXiv:2606. 13713v1 Announce Type: cross Abstract: Predicting cellular transcriptional responses to genetic perturbations is a central problem in single-cell biology, especially in the zero-shot setting where the perturbed gene or gene combination is unseen during training.
By Wei Zhang, Xun Jiang, Yuesi Xi, Ming Tang