AlphaGenome: AI for better understanding the genome
Introducing a new, unifying DNA sequence model that advances regulatory variant-effect prediction and promises to shed new light on genome function — now available via API.
AlphaGenome Atlas is a comprehensive resource that maps the molecular effects of 9 billion single‑letter DNA variants across the human genome. It provides a predictive map of how every possible DNA letter change could influence biological function. The atlas offers researchers a detailed view of variant impacts at an unprecedented scale.
Introducing a new, unifying DNA sequence model that advances regulatory variant-effect prediction and promises to shed new light on genome function — now available via API.
arXiv:2509.20702v3 Announce Type: replace-cross Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
arXiv:2605.21617v3 Announce Type: replace Abstract: Locating genomic features from genomic contact maps, such as centromere identification from genome-wide chromosome conformation capture techniques,...
arXiv:2511. 09026v2 Announce Type: replace-cross Abstract: Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood.
arXiv:2609.07500v1 Announce Type: cross Abstract: The evolution of DNA sequences can be viewed as stochastic dynamics on a high-dimensional discrete space, but it is unclear when empirical transition...
The study introduces WTKO-CNN, a convolutional neural network with an attention mechanism, to classify DNA sequences as wild‑type (WT) or knockout (KO) based on ATAC‑seq data. By generating saliency maps, the authors pinpointed influential nucleotide positions, extracted high‑saliency k‑mers, and performed de novo motif discovery, producing sequence logos and consensus motifs that align with known transcription factor binding sites. Validation with MEME, TOMTOM, and HOMER confirmed that the identified motifs belong to transcription factor families that differentiate WT from KO sequences.
EvoLen is a tokenization method for DNA language models that incorporates evolutionary information to prioritize functional sequence patterns such as regulatory motifs. It groups DNA sequences by cross-species evolutionary signals, trains separate BPE tokenizers for each group, merges vocabularies with a rule that favors preserved patterns, and uses length-aware decoding with dynamic programming. Experiments show EvoLen better preserves functional motifs, differentiates genomic contexts, and aligns with evolutionary constraints while matching or surpassing standard BPE on various DNALM benchmarks.
Clare Bryant uses Co-Scientist to identify genetic triggers in emerging infectious diseases.
arXiv:2607. 04987v1 Announce Type: new Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology.
Geneticist Catherine Brownstein demonstrates how OpenAI o1 can speed up the process of diagnosing rare medical challenges.
The paper introduces SIRGE, a sequence-informed geometric evaluator for RNA 3D structures that integrates nucleotide embeddings from a pretrained RNA language model into structural representations. SIRGE demonstrates superior performance over existing evaluators in Kendall–τ alignment, Top‑1 selection, and Top‑3 ranking. Controlled experiments reveal that sequence conditioning corrects errors of a purely geometric model and enhances target‑level ranking, suggesting that pretrained sequence representations provide complementary ranking information to geometric reasoning.
arXiv:2601. 12805v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks.