arXiv AI

LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation

LapDDPM is a conditional Graph Diffusion Probabilistic Model that generates high‑fidelity, biologically plausible single‑cell RNA sequencing data. It incorporates graph‑based inductive biases and a spectral adversarial perturbation mechanism to enforce robustness against structural noise, effectively acting as a Distributionally Robust Optimization framework. The model extends to spatial transcriptomics and multi‑modal data, and experimental results on datasets such as PBMC3K, Dentate Gyrus, HLCA, Visium, and 10x Multiome show it outperforms state‑of‑the‑art baselines in distribution matching, manifold preservation, and downstream utility.

arXiv AI
Aug 19

DMT-Dens: Density-preserving manifold visualization for biological data

DMT‑Dens is a parametric manifold‑visualization technique that uses a latent‑token Transformer encoder to produce two‑dimensional embeddings of high‑dimensional biological data. It preserves sampling density by aligning rank‑based manifold structures and optimizing a Pearson‑correlation loss on k‑nearest‑neighbor log‑radius estimates. Benchmark tests show that DMT‑Dens maintains density fidelity while achieving competitive label separability on biological datasets.

By Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang
arXiv Machine Learning
Aug 18

A Unified Geometric Framework for Developmental Analysis of Spatial Transcriptomic Data

arXiv:2608. 15306v1 Announce Type: cross Abstract: High-throughput single-cell and spatial transcriptomic technologies provide high-resolution snapshots of heterogeneous cellular states, but their destructive nature prevents repeated measurements of the same cells over time.

By Mary Chriselda Antony Oliver, Kaitlyn Hohmeier, Tuyen Tran, Alejandra Castillo, Caroline Moosm\"uller, Shiying Li
arXiv Machine Learning
Jun 5

HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data

arXiv:2506. 11152v4 Announce Type: replace-cross Abstract: Single-cell transcriptomics and proteomics have become a great source for data-driven insights into biology, enabling the use of advanced deep learning methods to understand cellular heterogeneity and gene expression at the single-cell level.

By Hiren Madhu, Jo\~ao Felipe Rocha, Tinglin Huang, Siddharth Viswanath, Smita Krishnaswamy, Rex Ying
arXiv AI
Jun 18

scGTN: Deep Siamese Graph Transformer Network for Single-cell RNA Sequencing Clustering

arXiv:2606. 18672v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) serves a pivotal role in characterizing gene expression at the cellular level, enabling the identification of cell types and advancing the understanding of cellular heterogeneity.

By Jinke Wu, Yifan Wang, Siyu Yi, Caiyang Yu, Ziyue Qiao, Nan Yin, Jiancheng Lv, Wei Ju
arXiv Machine Learning
Sep 25

SpaFactor: Lightweight Spatial Context-Aware Gene Program Modeling for Histology-to-Transcriptomics Inference

SpaFactor is a lightweight framework that predicts spatial gene expression from hematoxylin and eosin images by fusing central spot visuals with multiscale neighborhood context. It uses a residual MLP to map tissue microenvironment to low‑dimensional latent gene programs, which are decoded into coordinated multi‑gene predictions. Across five public cohorts, SpaFactor outperforms existing methods, especially for spatially variable genes, and better recovers biologically organized spatial patterns.

By Shiting Ruan, Xitong Ling, Qiming He, Ziyou Yan, Huaitian Yuan, Tian Guan, Ying Xiao, Xu Guan, Yonghong He
arXiv AI
3d ago

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.

By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie