arXiv:2603. 14717v2 Announce Type: replace Abstract: Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples.
By Jeffrey D. Varner
The paper introduces GenDA, a bidirectional discrete diffusion model designed for genomic sequence reconstruction, hypothesizing that entropy-guided span placement would improve variant-effect prediction and functional sequence generation. While the 202‑million‑parameter GenDA model achieves a higher ClinVar SNV AUROC (0.774) than a comparable autoregressive model, the improvement is not attributable to entropy guidance, and the model fails to outperform a shuffled‑gap baseline in zero‑shot functional inpainting across various genomic regions. The authors identify limitations such as tokenization granularity, span length caps, and the mismatch between local sequence complexity and functional importance, concluding that variant prediction, corruption priors, and functional generation are distinct tasks requiring separate validation.
By Susu Hu, Preetam Gattogi, Jens Lehmann, Sahar Vahdati, Stefanie Speidel, Julien Vibert
arXiv:2606. 18703v1 Announce Type: new Abstract: Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation.
By Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, Mar\'ia Rodr\'iguez Mart\'inez
arXiv:2609.07500v1 Announce Type: cross
Abstract: The evolution of DNA sequences can be viewed as stochastic dynamics on a high-dimensional discrete space, but it is unclear when empirical transition...
By Isabella Caranzano, Daniel Maria Busiello, Stefano Priorelli, Amos Maritan, Piero Fariselli
arXiv:2609.39644v2 Announce Type: new
Abstract: Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions...
By Gabriele Martino, Denis Skibinski, Ivo L. Hofacker, Sebastian Tschiatschek
arXiv:2606. 28659v1 Announce Type: cross Abstract: High-fidelity molecular docking simulations can produce biologically relevant estimates of epitope-receptor binding affinity but are computationally expensive and therefore limit the number of candidates that can be screened for vaccine design.
By Aspen Erlandsson Brisebois, Zahed Khatooni, Connor Burbridge, Brook Byrns, Heather L. Wilson, Sureesh Tikoo, Steven Rayan, Gordon Broderick