arXiv Machine Learning

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.

Hugging Face Trending Papers
Aug 19

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.

arXiv Machine Learning
Jun 9

Constraint-Aware Optimization for Robust Protein Stability Prediction

arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.

By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
arXiv AI
3d ago

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

SymFold introduces a symmetric dual‑path architecture that combines protein language models (PLMs) and multimodal protein language models (MPLMs) to iteratively guide protein sequence generation for inverse folding. By leveraging pretrained sequence evolution knowledge from PLMs and structural knowledge from MPLMs, the method improves upon the traditional serial pipeline where structure encoders produce coarse sequences refined by PLMs. Experiments on standard inverse‑folding benchmarks show state‑of‑the‑art performance, and ablation studies confirm the effectiveness of the symmetric design.

By Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
Hugging Face Trending Papers
Jun 17

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization

Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets. To overcome this supervision bottleneck, we introduce unsupervised reward optimization of PLMs, a comprehensive framework for steerable protein generation without ground-truth labels.

arXiv AI
Jun 2

Learning Implicit Bias in Generative Spaces for Accelerating Protein Dynamics Emulation

arXiv:2606. 01833v1 Announce Type: cross Abstract: Generative emulators of protein dynamics produce plausible trajectories at a fraction of the cost of molecular dynamics, but they inherit their training distribution and tend to revisit known states rather than reach rare ones under long-horizon extrapolation.

By Kaihui Cheng, Zhiqiang Cai, Wenkai Xiang, Zhihang Hu, Siyu Zhu, Tzuhsiung Yang, Yuan Qi
arXiv AI
Jun 9

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

arXiv:2507. 08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm.

By Changze Lv, Jiang Zhou, Siyu Long, Lihao Wang, Jiangtao Feng, Dongyu Xue, Yu Pei, Hao Wang, Zherui Zhang, Yuchen Cai, Zhiqiang Gao, Ziyuan Ma, Jiakai Hu, Chaochen Gao, Jingjing Gong, Yuxuan Song, Shuyi Zhang, Xiaoqing Zheng, Deyi Xiong, Lei Bai, Wanli Ouyang, Ya-Qin Zhang, Wei-Ying Ma, Bowen Zhou, Hao Zhou
arXiv Machine Learning
Aug 27

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

The paper introduces a method that uses orthogonal projection to remove the influence of known biochemical features from protein language model (PLM) embeddings, allowing the authors to assess how much these features contribute to protein fitness predictions. By applying this technique to high‑order and interaction effects, they demonstrate that eliminating these interpretable features reduces downstream classifier performance, indicating that PLM embeddings encode patterns correlated with biochemical properties. The authors also show that these biochemical features explain a substantial portion of the variance in the classifier’s predictions, suggesting that PLM embeddings capture biologically relevant information.

By Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard
arXiv AI
Aug 12

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

arXiv:2605. 02937v2 Announce Type: replace-cross Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity.

By Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya, Masashi Sugiyama, Li Erran Li, Jure Leskovec, Yejin Choi