Towards A Generative Protein Evolution Machine with DPLM-Evo
arXiv:2605. 00182v3 Announce Type: replace Abstract: Proteins are shaped by gradual evolution under biophysical and functional constraints.
The paper introduces GenDA, a bidirectional discrete diffusion model designed for genomic sequence reconstruction, hypothesizing that entropy-guided span placement would improve variant-effect prediction and functional sequence generation. While the 202‑million‑parameter GenDA model achieves a higher ClinVar SNV AUROC (0.774) than a comparable autoregressive model, the improvement is not attributable to entropy guidance, and the model fails to outperform a shuffled‑gap baseline in zero‑shot functional inpainting across various genomic regions. The authors identify limitations such as tokenization granularity, span length caps, and the mismatch between local sequence complexity and functional importance, concluding that variant prediction, corruption priors, and functional generation are distinct tasks requiring separate validation.
arXiv:2605. 00182v3 Announce Type: replace Abstract: Proteins are shaped by gradual evolution under biophysical and functional constraints.
arXiv:2608. 00697v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design.
arXiv:2607. 09039v1 Announce Type: new Abstract: The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability.
arXiv:2509. 26405v2 Announce Type: replace Abstract: We introduce InVirtuoGen, a discrete flow generative model for fragmented SMILES for de novo and fragment-constrained generation, and target-property/lead optimization of small molecules.
arXiv:2603. 14717v2 Announce Type: replace Abstract: Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples.
arXiv:2607. 29043v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important.
The study investigates whether unusually high‑gain parameters in transformer models—specifically gated feed‑forward network (gated‑FFN) rows—play a critical functional role across both text and genomic foundation models. By computing exact bilinear weight operators and testing structural extremeness, the authors find that high‑gain rows are enriched for functional importance but do not reliably predict causal effect size or severity. The analysis reveals model‑specific causal organizations, including super‑additive interactions in DNABERT‑2 and position‑localized dependencies in GENERator, indicating that structural prominence signals enrichment rather than calibrated criticality.
arXiv:2608.22849v2 Announce Type: replace Abstract: Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-...
arXiv:2606. 16044v1 Announce Type: new Abstract: Protein language models (pLMs) can generate novel protein sequences with properties beyond those observed in nature, yet the mechanisms underlying protein generation remain poorly understood.
arXiv:2606. 08802v1 Announce Type: new Abstract: Standard flow and diffusion pre-training matches the distribution of available data (e.
arXiv:2608. 14293v1 Announce Type: cross Abstract: High-content microscopy enables systematic profiling of cellular responses to chemical perturbations, but the scale of the chemical space makes exhaustive phenotypic characterization experimentally infeasible.
The paper introduces ORBIT, a framework for probing higher‑order epistasis in protein representations. ORBIT validates Walsh‑based diagnostics on synthetic landscapes, then applies them to the GB1 fitness landscape, comparing several models including ridge regression, MLPs, and Residual Interaction Tokenization (RIT). While no architecture differences were found in overall prediction performance, RIT notably increased pairwise token‑level accessibility, and deeper MLPs improved higher‑order functional recovery, revealing representation‑level changes hidden by conventional metrics.