arXiv AI

Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression

Predictor-Guided Latent Space Codon Optimization (LSCO) transforms the discrete codon selection problem into a continuous optimization task by embedding sequences into the latent space of a pretrained mRNA language model, allowing gradient-based search. LSCO integrates a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior derived from a protein-to-codon back-translation model, and constrained decoding to preserve protein fidelity. In a real-world wet-lab antibody expression dataset, LSCO outperforms both simple frequency-based methods and modern deep generative baselines in predicted expression while maintaining appropriate biophysical properties.

arXiv Machine Learning
Sep 4

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.

By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista
arXiv Machine Learning
Aug 21

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.

By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
arXiv Machine Learning
6d ago

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

The paper introduces IDiom, an autoregressive protein language model trained on a large dataset of intrinsically disordered protein regions (IDRs) from AlphaFold, and demonstrates that it can generate sequences matching natural IDR composition, motifs, and disorder. It further presents RL‑SAE, a reinforcement learning approach that uses sparse autoencoder features to steer generation toward specific functional patterns, achieving high activation of targeted features and improved predicted subcellular localization and transcriptional activity. The combination of IDiom and RL‑SAE allows interpretable, composable IDR design by explicitly controlling function‑associated sequence features.

By Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff
arXiv Machine Learning
6d ago

Analysis of Quantized and Efficiently Adapted Protein Language Models

The study evaluates 4‑bit quantization and low‑rank adapter fine‑tuning (QLoRA) on several large protein language models, finding that many model‑task pairs retain over 90% of full fine‑tuning performance while achieving up to 90% GPU memory savings. QLoRA preserves early‑layer representations and induces task‑specific changes in later layers, closely resembling full fine‑tuning with smaller representational shifts. For generative models, 4‑bit quantization largely maintains structural and sequence‑level properties, though token‑level analysis reveals model‑dependent changes in autoregressive output distributions.

By Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee
arXiv AI
Sep 7

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

ProtLingo is a protein language modeling framework that enhances a pretrained single‑sequence Transformer backbone with conditional local memory and sparse expert routing. It maps residue representations into discrete codes, composes local windows into latent N‑gram addresses, and retrieves reusable residual signals for recurring sequence contexts. The model also converts selected feed‑forward blocks into sparse Mixture‑of‑Experts layers, allowing residue‑dependent computation while activating only a subset of parameters, achieving competitive performance on protein fitness prediction, FLIP benchmarks, and supervised contact prediction with a 150M‑parameter backbone.

By Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu