arXiv Machine Learning

Improving scoring functions for protein-protein docking with LambdaLoss

The paper introduces LambdaDockScore, a protein‑protein docking scoring function that leverages the LambdaLoss loss from Learning‑to‑Rank to improve pose ranking. By fine‑tuning the energy prediction head of DFMDock on 2.9 million decoy poses from the DIPS dataset, LambdaDockScore outperforms the state‑of‑the‑art EuDockScore on CAPRI benchmark targets, achieving higher accuracy in top‑1 and top‑5 predictions. It also enhances ranking for antibody‑antigen and protein‑protein complexes with extreme binding interface sizes.

arXiv Machine Learning
Jun 5

An accurate nucleic acid-small molecule docking framework via geometric deep learning with large-scale pretraining

arXiv:2606. 05198v1 Announce Type: cross Abstract: Nucleic acids are increasingly recognized as therapeutic targets beyond conventional protein-centered drug discovery, yet accurate and efficient docking of small molecules to nucleic acid structures remains challenging.

By Shi Li (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Xujun Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Mingquan Liu (Faculty of Health Sciences, University of Macau, Macau SAR, China), Hui Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Shuoying Jia (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Yu Kang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Tingjun Hou (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China), Peichen Pan (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China)
arXiv Machine Learning
1d ago

StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes

StabilityArc is a method that decodes protein sequence embeddings into generalizable stability landscapes. It uses a shared RoPE transformer to map frozen ESMC-600M residue representations into an Lx20 matrix of substitution effects, with a symmetric, contact-aware residual to predict epistasis. In extensive leave-one-protein-out tests on 134,794 ProteinGym variants, StabilityArc achieves a Spearman correlation of 0.7134, surpassing the best zero‑shot baseline, and further improves Kermut’s performance when used as a prior.

By Aaron L. Feller, Andrew D. Ellington, Claus O. Wilke
arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

Hugging Face Trending Papers
Jul 20

Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus scoring and ensemble methods improve mean accuracy but treat all predictions identically without interpretable confidence measures or uncertainty decomposition, ignoring the chemical context of each protein-ligand pair.