Introducing a new, unifying DNA sequence model that advances regulatory variant-effect prediction and promises to shed new light on genome function — now available via API.
arXiv:2509.20702v3 Announce Type: replace-cross
Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
By Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou, Zhun Deng, Haoyu Zhang, Xihao Li, Didong Li
arXiv:2605.21617v3 Announce Type: replace
Abstract: Locating genomic features from genomic contact maps, such as centromere identification from genome-wide chromosome conformation capture techniques,...
By Elo\"ise Touron, Pedro L. C. Rodrigues, Julyan Arbel, Nelle Varoquaux, Michael Arbel
arXiv:2511. 09026v2 Announce Type: replace-cross Abstract: Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood.
By Pratik Dutta, Matthew Obusan, Rekha Sathian, Max Chao, Pallavi Surana, Nimisha Papineni, Yanrong Ji, Zhihan Zhou, Han Liu, Alisa Yurovsky, Ramana V Davuluri
arXiv:2609.07500v1 Announce Type: cross
Abstract: The evolution of DNA sequences can be viewed as stochastic dynamics on a high-dimensional discrete space, but it is unclear when empirical transition...
By Isabella Caranzano, Daniel Maria Busiello, Stefano Priorelli, Amos Maritan, Piero Fariselli
The study introduces WTKO-CNN, a convolutional neural network with an attention mechanism, to classify DNA sequences as wild‑type (WT) or knockout (KO) based on ATAC‑seq data. By generating saliency maps, the authors pinpointed influential nucleotide positions, extracted high‑saliency k‑mers, and performed de novo motif discovery, producing sequence logos and consensus motifs that align with known transcription factor binding sites. Validation with MEME, TOMTOM, and HOMER confirmed that the identified motifs belong to transcription factor families that differentiate WT from KO sequences.
By Lopamudra Dey