arXiv Machine Learning By Aaron L. Feller, Andrew D. Ellington, Claus O. Wilke

StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes

Read the original on arXiv Machine Learning →

StabilityArc is a method that decodes protein sequence embeddings into generalizable stability landscapes. It uses a shared RoPE transformer to map frozen ESMC-600M residue representations into an Lx20 matrix of substitution effects, with a symmetric, contact-aware residual to predict epistasis. In extensive leave-one-protein-out tests on 134,794 ProteinGym variants, StabilityArc achieves a Spearman correlation of 0.7134, surpassing the best zero‑shot baseline, and further improves Kermut’s performance when used as a prior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 9

Constraint-Aware Optimization for Robust Protein Stability Prediction

arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.

By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
arXiv Machine Learning
Aug 27

Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations

The paper introduces ORBIT, a framework for probing higher‑order epistasis in protein representations. ORBIT validates Walsh‑based diagnostics on synthetic landscapes, then applies them to the GB1 fitness landscape, comparing several models including ridge regression, MLPs, and Residual Interaction Tokenization (RIT). While no architecture differences were found in overall prediction performance, RIT notably increased pairwise token‑level accessibility, and deeper MLPs improved higher‑order functional recovery, revealing representation‑level changes hidden by conventional metrics.

By Maryam Rahimimovassagh, Ivan Garibay, Niloofar Yousefi
arXiv Machine Learning
Sep 23

STAR-VAE: A Scalable Latent-Variable Transformer for Controllable Molecular Generation

STAR-VAE is a Transformer-based variational autoencoder that uses SELFIES encoding and a bidirectional encoder with an autoregressive decoder pretrained on 79 million PubChem molecules. It incorporates a property signal to jointly condition the prior, posterior, and decoder, and employs LoRA adapters for fine‑tuning on small datasets without altering the backbone. The model achieves 100 % validity and near‑perfect novelty in MOSES sampling, low KL divergence on several GuacaMol descriptors, strong synthetic‑accessibility conditioning, and effective docking‑score control across multiple protein targets, while also enabling scaffold recovery and diverse label‑conditioned generation on ChEMBL targets.

By Bum Chul Kwon, Ben Shapira, Moshiko Raboh, Shreyans Sethi, Shruti Murarka, Joseph A Morrone, Leili Zhang, Wendy Cornell, Jianying Hu, Parthasarathy Suryanarayanan
arXiv AI
Sep 25

PFArena: Benchmarking Language Models for Protein Modification

PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.

By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou
arXiv AI
Jun 9

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

arXiv:2507. 08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm.

By Changze Lv, Jiang Zhou, Siyu Long, Lihao Wang, Jiangtao Feng, Dongyu Xue, Yu Pei, Hao Wang, Zherui Zhang, Yuchen Cai, Zhiqiang Gao, Ziyuan Ma, Jiakai Hu, Chaochen Gao, Jingjing Gong, Yuxuan Song, Shuyi Zhang, Xiaoqing Zheng, Deyi Xiong, Lei Bai, Wanli Ouyang, Ya-Qin Zhang, Wei-Ying Ma, Bowen Zhou, Hao Zhou
arXiv Machine Learning
Jun 4

Structure-Aware Prediction of PROTAC-Mediated Protein Degradability via Graph Neural Networks

arXiv:2606. 04021v1 Announce Type: cross Abstract: Proteolysis-targeting chimeras (PROTACs) can selectively degrade disease-causing proteins, yet predicting which targets are amenable to degradation remains a critical bottleneck: existing computational methods require the complete PROTAC molecular structure, information unavailable before synthesis.

By Bryan Cheng, Austin Jin