The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.
By Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang
arXiv:2606. 18961v1 Announce Type: new Abstract: Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets.
By Lanqing Li, Shentong Mo, Yang Yu, Pheng-Ann Heng
SymFold introduces a symmetric dual‑path architecture that combines protein language models (PLMs) and multimodal protein language models (MPLMs) to iteratively guide protein sequence generation for inverse folding. By leveraging pretrained sequence evolution knowledge from PLMs and structural knowledge from MPLMs, the method improves upon the traditional serial pipeline where structure encoders produce coarse sequences refined by PLMs. Experiments on standard inverse‑folding benchmarks show state‑of‑the‑art performance, and ablation studies confirm the effectiveness of the symmetric design.
By Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets. To overcome this supervision bottleneck, we introduce unsupervised reward optimization of PLMs, a comprehensive framework for steerable protein generation without ground-truth labels.
arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
arXiv:2606. 02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints.
By Sahil Rahman, Maxx Richard Rahman