arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
arXiv:2609.37808v1 Announce Type: new
Abstract: Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific p...
By Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu
Q-BIOLAT is a framework that converts pretrained protein-language-model embeddings into compact binary codes and trains a quadratic unconstrained binary optimization (QUBO) surrogate with unary and pairwise latent interactions for protein fitness optimization. The study demonstrates that binary encodings with similar predictive accuracy can produce different Hamming neighborhoods, affecting local optima and search trajectories, and shows that PCA followed by per‑coordinate median thresholding yields a more balanced binary space than AE/VAE baselines. Experimental evaluation on GFP and AAV fitness landscapes from ProteinGym confirms that simulated annealing, genetic algorithms, and greedy hill climbing can retrieve high‑percentile variants, with decoded candidates reported via surrogate‑predicted scores.
By Truong-Son Hy
arXiv:2605. 01625v3 Announce Type: replace Abstract: Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial resolutions, ranging from atomic interactions to global fold topology.
By Viet Thanh Duy Nguyen, John K. Johnstone, Truong-Son Hy
arXiv:2608.29207v1 Announce Type: new
Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (t...
By Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao
arXiv:2605.16581v2 Announce Type: replace
Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual...
By Thomas Walton, Ayan Goel, Amirali Aghazadeh
arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.
By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
arXiv:2607. 20551v1 Announce Type: cross Abstract: Effective molecular representation learning is crucial for accurate molecular property prediction.
By Tianming Han, Li Zhang, Qi Zhao
arXiv:2604.18467v3 Announce Type: replace-cross
Abstract: Motivation: Peptide-protein interactions (PepPIs) are central to cellular regulation and peptide therapeutics, but experimental characterizat...
By Chupei Tang, Junxiao Kong, Moyu Tang, Di Wang, Jixiu Zhai, Ronghao Xie, Shangkun Sima, Tianchi Lu
arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.
By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
By Davy Guan, Lu Zhang, Asiri Wijesinghe, Allen Zhu, He Zhao, Helen Power, F. Hafna Ahmed, Andrew Warden, Cheng Soon Ong, Daniel M. Steinberg
arXiv:2608. 15483v1 Announce Type: new Abstract: Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers.
By Fanqi Wang, Weisheng Tang, Hairong Qi