arXiv:2608. 11269v1 Announce Type: cross Abstract: Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging.
By Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
We describe a systematic approach for spawning and aggregating multi-class cryo-EM reconstruction jobs. This approach formalizes standard ad hoc strategies of iterative classification and filtering ty...
The article presents a systematic meta-algorithm for spawning and aggregating multi-class cryo-EM reconstruction jobs, formalizing iterative classification and filtering strategies used by practitioners. It claims to be the first method capable of ab initio reconstruction on datasets with dozens of distinct species, achieving 97% accuracy on a 45-class subset of Tomotwin-100 and 75% on the full dataset, and successfully recovering ribosomal assembly states from an unfiltered experimental cryo-EM dataset. The approach scales with compute resources and aims to underpin automated cryo-EM workflows in contemporary experimental settings.
By Alkin Kaz, Arda Kaz, Ellen D. Zhong
arXiv:2609.38538v1 Announce Type: cross
Abstract: Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as...
By Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain
arXiv:2607. 04987v1 Announce Type: new Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology.
By Dmytro Rizdvanetskyi, Nathan Ross, Pavlo Lutsik
arXiv:2607. 23518v1 Announce Type: new Abstract: The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination.
By Hengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang, Yunhong Lu, Chenxi Cai, Yan Zhang, Min Zhang
arXiv:2602. 04901v2 Announce Type: replace-cross Abstract: Predicting transcriptional responses to genetic perturbations is a central problem in functional genomics.
By Jiafa Ruan, Ruijie Quan, Liyang Xu, Zongxin Yang, Yi Yang
SpaFactor is a lightweight framework that predicts spatial gene expression from hematoxylin and eosin images by fusing central spot visuals with multiscale neighborhood context. It uses a residual MLP to map tissue microenvironment to low‑dimensional latent gene programs, which are decoded into coordinated multi‑gene predictions. Across five public cohorts, SpaFactor outperforms existing methods, especially for spatially variable genes, and better recovers biologically organized spatial patterns.
By Shiting Ruan, Xitong Ling, Qiming He, Ziyou Yan, Huaitian Yuan, Tian Guan, Ying Xiao, Xu Guan, Yonghong He
The paper investigates adapting the HiFiC generative image compression model, originally designed for natural photographs, to compress Hi‑C chromatin contact maps while preserving biologically relevant features. By replacing HiFiC’s distortion term with a spatially‑weighted MSE that emphasizes loops, TAD boundaries, stripes, and compartments, and adding an insulation‑score loss, the authors fine‑tune a pretrained HiFiC checkpoint in a three‑phase strategy to create HiFiC‑G. Evaluation on two cell lines shows HiFiC‑G better retains local structures such as stripes and TAD boundaries, though long‑range A/B compartment preservation remains limited due to architectural constraints.
By Andre Antonio Straton
Compositional data -- vectors encoding relative proportions -- arise across scientific domains, including ecology, geochemistry, and genomics. The features in these data often come with known hierarchical structure (e.
WEECFP-SuRGE introduces a position‑aware substructure encoding method that combines tokenized hierarchical Morgan fingerprints with graph‑distance‑dependent rotations applied at the input and within transformer self‑attention. The approach captures local chemistry, long‑range interactions, and molecular topology without requiring external pretraining or 3‑D conformer generation. Benchmarks on MoleculeNet and the Therapeutic Data Commons ADMET datasets show competitive performance, and a reconstruction procedure correctly identifies constitutional isomers for 92.6% of a 4,200‑molecule library.
By Robert Epps
arXiv:2607. 22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations.
By Chen Wang, Boming Kang, Qinghua Cui