AlphaGenome Atlas is a comprehensive resource that maps the molecular effects of 9 billion single‑letter DNA variants across the human genome. It provides a predictive map of how every possible DNA letter change could influence biological function. The atlas offers researchers a detailed view of variant impacts at an unprecedented scale.
arXiv:2511. 09026v2 Announce Type: replace-cross Abstract: Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood.
By Pratik Dutta, Matthew Obusan, Rekha Sathian, Max Chao, Pallavi Surana, Nimisha Papineni, Yanrong Ji, Zhihan Zhou, Han Liu, Alisa Yurovsky, Ramana V Davuluri
EvoLen is a tokenization method for DNA language models that incorporates evolutionary information to prioritize functional sequence patterns such as regulatory motifs. It groups DNA sequences by cross-species evolutionary signals, trains separate BPE tokenizers for each group, merges vocabularies with a rule that favors preserved patterns, and uses length-aware decoding with dynamic programming. Experiments show EvoLen better preserves functional motifs, differentiates genomic contexts, and aligns with evolutionary constraints while matching or surpassing standard BPE on various DNALM benchmarks.
By Nan Huang, Xiaoxiao Zhou, Junxia Cui, Mario Tapia-Pacheco, Tiffany Amariuta, Yang Li, Jingbo Shang
arXiv:2607. 00931v1 Announce Type: new Abstract: Predicting cancer drug response from transcriptomic profiles is a cornerstone of precision oncology, yet the scientific value of machine learning models hinges not solely on predictive accuracy, but also on their capacity to generate reliable biological insights.
By Martino Ciaperoni, Margherita Lalli, Simone Piaggesi, Martina Varisco, Francesco Carli, Riccardo Guidotti, Dino Pedreschi, Francesco Raimondi, Fosca Giannotti
The study introduces WTKO-CNN, a convolutional neural network with an attention mechanism, to classify DNA sequences as wild‑type (WT) or knockout (KO) based on ATAC‑seq data. By generating saliency maps, the authors pinpointed influential nucleotide positions, extracted high‑saliency k‑mers, and performed de novo motif discovery, producing sequence logos and consensus motifs that align with known transcription factor binding sites. Validation with MEME, TOMTOM, and HOMER confirmed that the identified motifs belong to transcription factor families that differentiate WT from KO sequences.
By Lopamudra Dey
arXiv:2609.14882v1 Announce Type: cross
Abstract: Nucleotide sequence analysis is central to problems spanning regulatory genomics, evolutionary biology, and phenotype prediction. Classical bioinform...
By Evgeny S. Saveliev, Krzysztof Kacprzyk, Charlotte Capitanchik, Neelanjan Mukherjee, Kate Matlin, Ryan Sheridan, Srinivas Ramachandran, Jernej Ule, David L. Bentley, Mihaela van der Schaar
arXiv:2509.20702v3 Announce Type: replace-cross
Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
By Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou, Zhun Deng, Haoyu Zhang, Xihao Li, Didong Li
arXiv:2606. 14734v1 Announce Type: cross Abstract: Motivation: Gene regulatory network inference from single-cell RNA sequencing (scRNA-seq) data is important for uncovering cell-state-specific transcriptional programs.
By Ziyang Dong, Shanwen Tan, Hengchuang Yin, Wei Liu, Yifan Wang, Siyu Yi, Jiancheng Lv, Wei Ju
arXiv:2602. 04901v2 Announce Type: replace-cross Abstract: Predicting transcriptional responses to genetic perturbations is a central problem in functional genomics.
By Jiafa Ruan, Ruijie Quan, Liyang Xu, Zongxin Yang, Yi Yang
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
arXiv:2608. 14330v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables genome-wide gene expression profiling while preserving tissue architecture, but its cost and limited scalability remain major bottlenecks.
By Ruyter Swann, Dorent Reuben, Racoceanu Daniel
arXiv:2606. 02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather than merely fit static measurements.
By Jin Gao, Juntu Zhao, Zirui Zeng, Jiaqi Shen, Junhao Shi, Dukun Zhao, Yuming Lu, Dequan Wang