arXiv AI By Alberto Caron, Tianyu Cui, Dmytro S. Lituiev, Mangal Prakash, Artem Moskalev, Amina Mollaysa, Bo Zhai, Hirsh Nanda, Daniel M. Poole, Zhongyin Liu, Iman Farasat, Robert Davidson, Nikolay V. Manyakov, Tommaso Mansi, Scott Oloff, Rui Liao

Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression

Read the original on arXiv AI →

Predictor-Guided Latent Space Codon Optimization (LSCO) transforms the discrete codon selection problem into a continuous optimization task by embedding sequences into the latent space of a pretrained mRNA language model, allowing gradient-based search. LSCO integrates a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior derived from a protein-to-codon back-translation model, and constrained decoding to preserve protein fidelity. In a real-world wet-lab antibody expression dataset, LSCO outperforms both simple frequency-based methods and modern deep generative baselines in predicted expression while maintaining appropriate biophysical properties.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.

By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista
arXiv Machine Learning
Aug 21

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.

By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
arXiv AI
Aug 12

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".

By Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
arXiv Machine Learning
6d ago

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

The paper introduces IDiom, an autoregressive protein language model trained on a large dataset of intrinsically disordered protein regions (IDRs) from AlphaFold, and demonstrates that it can generate sequences matching natural IDR composition, motifs, and disorder. It further presents RL‑SAE, a reinforcement learning approach that uses sparse autoencoder features to steer generation toward specific functional patterns, achieving high activation of targeted features and improved predicted subcellular localization and transcriptional activity. The combination of IDiom and RL‑SAE allows interpretable, composable IDR design by explicitly controlling function‑associated sequence features.

By Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff