arXiv Machine Learning

Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks

arXiv:2602. 06020v3 Announce Type: replace Abstract: How do protein structure prediction models fold proteins?

arXiv AI
Sep 2

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

SymFold introduces a symmetric dual‑path architecture that combines protein language models (PLMs) and multimodal protein language models (MPLMs) to iteratively guide protein sequence generation for inverse folding. By leveraging pretrained sequence evolution knowledge from PLMs and structural knowledge from MPLMs, the method improves upon the traditional serial pipeline where structure encoders produce coarse sequences refined by PLMs. Experiments on standard inverse‑folding benchmarks show state‑of‑the‑art performance, and ablation studies confirm the effectiveness of the symmetric design.

By Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
arXiv Machine Learning
Jun 29

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

arXiv:2606. 27440v1 Announce Type: new Abstract: Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules.

By Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts, Nicholas Franklin, Axel Elaldi, Olivia Viessmann
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv AI
Aug 19

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

The paper introduces MCTH (Monte Carlo Tree Hallucination), an inference-only framework that performs all‑atom biomolecular sequence‑structure co‑design by treating pretrained folding and inverse‑folding models as black‑box operators. MCTH uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence, uncertainty, and cross‑expert consensus. Experiments across protein‑RNA, protein‑DNA, protein‑protein, and protein‑ligand design show that adaptive search outperforms simpler sampling strategies, and evaluations with AlphaFold3 and Chai‑1 demonstrate transferability beyond the search‑time oracle.

By Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens
arXiv Machine Learning
Jun 16

Learning Topological Representations for Molecular Dynamics

arXiv:2606. 14737v1 Announce Type: cross Abstract: Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted observables or learned kinetic embeddings.

By Dominik Geng, Florian Graf, Martin Uray, Roland Kwitt