arXiv:2605.07938v2 Announce Type: replace
Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular funct...
By Sachini Weerasekara, Natasha Darras, Sagar Kamarthi, Colles Price, Jacqueline Isaacs
arXiv:2512. 17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling.
By Daphn\'e Chopard, Jorge da Silva Gon\c{c}alves, Irene Cannistraci, Thomas M. Sutter, Julia E. Vogt
arXiv:2608. 00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values.
By Jiaqi Xiong, Yuntao hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information using a cross‑attention architecture. This approach models interactions within distinct subcellular compartments, producing fine‑grained embeddings that capture both molecular expression patterns and functional protein properties. It is presented as the first method to jointly incorporate transcriptomic data, sequence, and structural knowledge for subcellularly resolved cell representation.
By Zhen Zhou, Jiachen Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen
arXiv:2606. 00685v1 Announce Type: new Abstract: Gene regulatory networks (GRNs) capture transcription factor-target interactions and are central to understanding cell-state regulation and disease.
By Tianyang Xu, Tianci Liu, Niraj Rayamajhi, Ryan Patrick, Kranthi Varala, Ying Li, Jing Gao
The paper introduces scTrilemma, a latent-bottleneck variational autoencoder designed to address the representation trilemma in single‑cell RNA‑seq data: preserving biological identity and state, remaining robust to nuisance context, and retaining gene‑level variation for expression analysis. scTrilemma routes expression‑derived variation to the embedding, decoder, or prior, gating gene tokens by expression and conditioning the prior on unlabeled pseudo‑bulk context, all under a single reconstruction objective without target annotations. In zero‑shot evaluations on successive CZ CELLxGENE Census releases, scTrilemma simultaneously satisfies all three demands, maintaining biological state, differential‑expression, and pathway structure across multiple disease settings, and latent interventions show context can be removed with minimal impact on other demands.
By Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, Seonghwan Kim, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung Park
The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.
By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca
The paper introduces a multimodal framework that learns subcellularly resolved cell embeddings by integrating RNA expression profiles, protein sequence representations, and protein structural information. It uses a cross‑attention architecture to model interactions across distinct subcellular compartments, producing embeddings that capture both molecular expression patterns and functional protein properties. This approach is presented as the first to jointly incorporate transcriptomic data, protein sequences, and structural knowledge within a unified cross‑modal learning paradigm.
arXiv:2609.36429v1 Announce Type: new
Abstract: Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods...
By Zijun Gao, Chunbin Gu, Jinxi Xiang, Xiangde Luo, Pheng-Ann Heng
arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.
By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
arXiv:2608. 05928v1 Announce Type: new Abstract: Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes.
By Yuhao Wang, Zelin Zang, Yuxuan Liu, Zhen Lei, Stan Z. Li
arXiv:2506. 11152v4 Announce Type: replace-cross Abstract: Single-cell transcriptomics and proteomics have become a great source for data-driven insights into biology, enabling the use of advanced deep learning methods to understand cellular heterogeneity and gene expression at the single-cell level.
By Hiren Madhu, Jo\~ao Felipe Rocha, Tinglin Huang, Siddharth Viswanath, Smita Krishnaswamy, Rex Ying