arXiv AI
2d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv AI
Sep 15

Towards a knowledge-enhanced single-cell foundation model

The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.

By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca
arXiv Machine Learning
Jun 9

Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

arXiv:2606. 09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supervised learning on millions of cells.

By Mikele Milia, Louis Fabrice Tshimanga, Henning Mueller, Manfredo Atzori, Barbara Di Camillo
arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung