arXiv Machine Learning

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

Hugging Face Trending Papers
Aug 19

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.

arXiv Machine Learning
Aug 20

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.

By Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang
arXiv Machine Learning
Jun 10

Flexible Kernels for Protein Property Prediction

arXiv:2606. 11057v1 Announce Type: new Abstract: Despite its importance to applications in protein design, predicting protein properties like binding affinity and thermostability from sparse experimental data remains a significant challenge.

By Martin Jankowiak, Yerdos Ordabayev, Rudraksh Tuwani, Henry N. Ward, Hunter Nisonoff, James M. McFarland, Gevorg Grigoryan
arXiv AI
Sep 7

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

ProtLingo is a protein language modeling framework that enhances a pretrained single‑sequence Transformer backbone with conditional local memory and sparse expert routing. It maps residue representations into discrete codes, composes local windows into latent N‑gram addresses, and retrieves reusable residual signals for recurring sequence contexts. The model also converts selected feed‑forward blocks into sparse Mixture‑of‑Experts layers, allowing residue‑dependent computation while activating only a subset of parameters, achieving competitive performance on protein fitness prediction, FLIP benchmarks, and supervised contact prediction with a 150M‑parameter backbone.

By Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu
arXiv Machine Learning
Sep 4

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.

By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista