arXiv:2607. 15018v1 Announce Type: cross Abstract: High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables.
By Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, ShengLi Tzeng, Han-Ming Wu
DMT‑Dens is a parametric manifold‑visualization technique that uses a latent‑token Transformer encoder to produce two‑dimensional embeddings of high‑dimensional biological data. It preserves sampling density by aligning rank‑based manifold structures and optimizing a Pearson‑correlation loss on k‑nearest‑neighbor log‑radius estimates. Benchmark tests show that DMT‑Dens maintains density fidelity while achieving competitive label separability on biological datasets.
By Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang
arXiv:2604.02535v2 Announce Type: replace
Abstract: Dimensionality reduction (DR) involves two longstanding trade-offs. First, preserving local neighborhoods can come at the cost of global structure....
By Zeyang Huang, Angelos Chatzimparmpas, Thomas H\"ollt, Takanori Fujiwara
SpaFactor is a lightweight framework that predicts spatial gene expression from hematoxylin and eosin images by fusing central spot visuals with multiscale neighborhood context. It uses a residual MLP to map tissue microenvironment to low‑dimensional latent gene programs, which are decoded into coordinated multi‑gene predictions. Across five public cohorts, SpaFactor outperforms existing methods, especially for spatially variable genes, and better recovers biologically organized spatial patterns.
By Shiting Ruan, Xitong Ling, Qiming He, Ziyou Yan, Huaitian Yuan, Tian Guan, Ying Xiao, Xu Guan, Yonghong He
arXiv:2506. 11152v4 Announce Type: replace-cross Abstract: Single-cell transcriptomics and proteomics have become a great source for data-driven insights into biology, enabling the use of advanced deep learning methods to understand cellular heterogeneity and gene expression at the single-cell level.
By Hiren Madhu, Jo\~ao Felipe Rocha, Tinglin Huang, Siddharth Viswanath, Smita Krishnaswamy, Rex Ying
arXiv:2607. 14410v1 Announce Type: new Abstract: Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines.
By Jagan Mohan Reddy Dwarampudi, Veena Kochat, Suresh Satpati, Kunal Rai, Tania Banerjee
arXiv:2506. 22228v2 Announce Type: replace-cross Abstract: Single-cell sequencing is revolutionizing biology by enabling detailed investigations of cell-state transitions.
By Rong Ma, Xi Li, Jingyuan Hu, Bin Yu
arXiv:2605. 23540v2 Announce Type: replace Abstract: Dimensionality Reduction (DR) methods are widely used to visualize high-dimensional data.
By Diede P. M. van der Hoorn, Alessio Arleo, Fernando V. Paulovich
arXiv:2606. 13007v1 Announce Type: cross Abstract: Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity.
By Ping Xu, Pengjiang Li, Tian Du, Zaitian Wang, Jiawei Gu, Ziyue Qiao, Pengfei Wang, Yuanchun Zhou
The paper introduces the Rashomon set for dimension reduction, a collection of equally good embeddings that preserve high‑dimensional structure. It proposes PCA‑informed alignment to make axes interpretable, concept‑alignment regularization to incorporate external knowledge, and a method to extract trustworthy nearest‑neighbor relationships across the Rashomon set for refined embeddings. These techniques aim to produce interpretable, robust, and goal‑aligned visualizations by leveraging multiple valid embeddings instead of a single one.
By Yiyang Sun, Haiyang Huang, Gaurav Rajesh Parikh, Cynthia Rudin
LapDDPM is a conditional Graph Diffusion Probabilistic Model that generates high‑fidelity, biologically plausible single‑cell RNA sequencing data. It incorporates graph‑based inductive biases and a spectral adversarial perturbation mechanism to enforce robustness against structural noise, effectively acting as a Distributionally Robust Optimization framework. The model extends to spatial transcriptomics and multi‑modal data, and experimental results on datasets such as PBMC3K, Dentate Gyrus, HLCA, Visium, and 10x Multiome show it outperforms state‑of‑the‑art baselines in distribution matching, manifold preservation, and downstream utility.
By Lorenzo Bini, Stephane Marchand-Maillet
arXiv:2606. 31394v1 Announce Type: cross Abstract: Artificial intelligence is transforming our capability to solve biological challenges.
By Jisung Park, Seohyeon Kang, Daeun Yoo, Eunsu Lee, Seoin Cho, Wooyeop Choi, Ian Choi, James R. Evan, Daesoo Kim, Sonia Gandhi, Minee L. Choi