MIT News AI

Looking beyond natural sequences

A new machine‑learning framework is being developed to enhance the success rate of computational protein design. It deliberately moves away from reproducing sequences found in nature, aiming instead for novel designs that may perform better in practical applications.

Hugging Face Trending Papers
Sep 2

Advances in Machine Learning for Directed Evolution: A Five-Year Retrospective

The article reviews the past five years of machine learning (ML) applications in protein engineering, noting that directed evolution has not benefited as much as other disciplines. It argues that a mismatch between ML‑assisted directed evolution (MLDE) goals—finding an optimal protein—and broader directed evolution aims—finding a sufficient protein within time and resource limits—has hindered progress. The author points out that most MLDE methods ignore DNA synthesis costs, limiting practical use, and concludes by highlighting recent exceptions and suggesting that MLDE objectives can be reframed to align with real‑world constraints.

arXiv Machine Learning
Sep 4

Advances in Machine Learning for Directed Evolution: A Five-Year Retrospective

The article reviews the past five years of machine learning (ML) applications in protein engineering, noting that directed evolution has not benefited as much as other disciplines. It argues that a mismatch between ML‑assisted directed evolution (MLDE) goals—finding an optimal protein—and broader directed evolution aims—finding a sufficient protein within time and resource limits—has hindered progress. The author points out that most MLDE methods ignore DNA synthesis costs, limiting practical use, and concludes by highlighting recent exceptions and suggesting that MLDE objectives can be reframed to align with real‑world constraints.

By Bruce J. Wittmann
arXiv AI
Jun 9

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

arXiv:2507. 08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm.

By Changze Lv, Jiang Zhou, Siyu Long, Lihao Wang, Jiangtao Feng, Dongyu Xue, Yu Pei, Hao Wang, Zherui Zhang, Yuchen Cai, Zhiqiang Gao, Ziyuan Ma, Jiakai Hu, Chaochen Gao, Jingjing Gong, Yuxuan Song, Shuyi Zhang, Xiaoqing Zheng, Deyi Xiong, Lei Bai, Wanli Ouyang, Ya-Qin Zhang, Wei-Ying Ma, Bowen Zhou, Hao Zhou
Hugging Face Trending Papers
Aug 19

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.

arXiv Machine Learning
Aug 20

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.

By Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang
arXiv AI
3d ago

Structure-aware Reinforcement Learning for Protein Directed Evolution

The paper introduces StructEvo, a structure-aware reinforcement learning framework designed to improve protein directed evolution. By using a delta-structure fusion encoder to approximate mutant structure features and a hierarchical action network aligned with protein structure, the method navigates the vast mutation space more effectively. StructEvo outperforms existing machine learning-assisted directed evolution techniques by 9.2% and 16.3% on two benchmarks and uncovers an experimentally validated epistasis pattern in GFP, underscoring the value of structural guidance.

By Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
arXiv Machine Learning
Jun 10

Flexible Kernels for Protein Property Prediction

arXiv:2606. 11057v1 Announce Type: new Abstract: Despite its importance to applications in protein design, predicting protein properties like binding affinity and thermostability from sparse experimental data remains a significant challenge.

By Martin Jankowiak, Yerdos Ordabayev, Rudraksh Tuwani, Henry N. Ward, Hunter Nisonoff, James M. McFarland, Gevorg Grigoryan
arXiv AI
Sep 2

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

SymFold introduces a symmetric dual‑path architecture that combines protein language models (PLMs) and multimodal protein language models (MPLMs) to iteratively guide protein sequence generation for inverse folding. By leveraging pretrained sequence evolution knowledge from PLMs and structural knowledge from MPLMs, the method improves upon the traditional serial pipeline where structure encoders produce coarse sequences refined by PLMs. Experiments on standard inverse‑folding benchmarks show state‑of‑the‑art performance, and ablation studies confirm the effectiveness of the symmetric design.

By Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
arXiv Machine Learning
1d ago

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

The paper introduces IDiom, an autoregressive protein language model trained on a large dataset of intrinsically disordered protein regions (IDRs) from AlphaFold, and demonstrates that it can generate sequences matching natural IDR composition, motifs, and disorder. It further presents RL‑SAE, a reinforcement learning approach that uses sparse autoencoder features to steer generation toward specific functional patterns, achieving high activation of targeted features and improved predicted subcellular localization and transcriptional activity. The combination of IDiom and RL‑SAE allows interpretable, composable IDR design by explicitly controlling function‑associated sequence features.

By Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff