arXiv:2607. 17601v1 Announce Type: cross Abstract: Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust.
By Yongchan Hong, Defu Cao, Wenjin Liu, Thomas Ku, Jordy Homing Lam, Emily Nguyen, Willie Neiswanger, Vsevolod Katritch, Yan Liu
CaliPPer is a post‑hoc framework that calibrates and predicts the performance of binding‑prediction models by combining a multi‑chain Sample‑to‑Domain Distance (S2DD) metric with distance‑aware Bayesian recalibration. It operates at three resolutions—generalisability score, aggregate performance prediction, and per‑sample confidence—achieving strong distance‑performance correlations (|r| = 0.80–0.92) and low prediction errors for AUROC, AP, and F1. In retrospective analyses of five published studies, CaliPPer increased true discovery rates, improving AUROC by up to +0.20 on unseen epitopes and variants and raising confirmed neoantigen findings from 0/5 to 3/5.
By Jian-Qing Zheng, Hantao Lou, Zinan Yin, Sam Farrar, Yuze Zhou, Elie Antoun, Xiangxi Wang, Xuetao Cao, Tao Dong
ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.
ProbeMatchDTI is a new framework for drug‑target interaction prediction that uses probe‑driven pattern matching to preserve weak biochemical signals. It introduces IterProbe, which retains contextual states across refinement depths and selects them with learnable probes, and BindingProbe, which models drug‑protein complementarity at both local and whole‑pair levels. Experiments show that ProbeMatchDTI outperforms existing methods, improving AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank, and its predictions can be integrated into downstream drug‑discovery workflows.
By Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
The paper introduces a conformal prediction framework designed for molecular property prediction under label shift. By weighting conformal scores with marginal label probability ratios, it generates statistically rigorous prediction intervals without retraining, enabling robust uncertainty quantification when property distributions change. This approach provides actionable confidence measures that improve the reliability of AI-driven predictions in drug discovery.
By Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin
arXiv:2603. 10950v2 Announce Type: replace Abstract: Machine learning methods for identifying molecular structures from tandem mass spectra (MS/MS) have advanced rapidly, yet current approaches still exhibit significant error rates.
By Mira J\"urgens, Gaetan De Waele, Morteza Rakhshaninejad, Willem Waegeman
arXiv:2607. 19237v1 Announce Type: new Abstract: Designing small molecule ligands that bind with high affinity to specific protein pockets is a fundamental goal in drug discovery, as small molecules constitute a major fraction of approved therapeutics.
By Yiming Qin, Kai Yi, Miruna Cretu, Sjors H. W. Scheres, Pietro Li\`o, Pascal Frossard
arXiv:2608. 09099v1 Announce Type: new Abstract: Quantitative estimation of protein-ligand binding affinity from three-dimensional complex structures is a fundamental task in structure-based computational chemistry and molecular modeling.
By Qingyang Zou, Jiaye Huang, Hangbo Xie, Jiayue Yin, Youyi Song, Jinfeng Liu
arXiv:2607. 11091v1 Announce Type: new Abstract: A trained molecular property model can be refined at test time by correcting each prediction with the measured labels of the most similar training molecules, a retraining-free procedure we call neighbor fusion; evidential neural networks make it principled by using their aleatoric and epistemic uncertainty to parameterize a Bayesian update.
By Cameron Gruich, Weichi Yao, Yixin Wang, Bryan Goldsmith
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
The paper introduces ReGeoDTA, a framework that preserves chemical heterogeneity and continuous geometric relationships in drug and protein representations to improve drug–target affinity prediction. Experiments on three benchmark datasets show that maintaining representation fidelity consistently enhances predictive accuracy across various DTA architectures, while degrading representations harms performance and cannot be recovered by more complex downstream models. The study highlights representation fidelity as a key upstream design principle for accurate and generalizable affinity prediction.
By Yixiao Li, Yining Qian, Yefan Chen, Zenghui Chen, Jiayue Sun, Yuhai Zhao, Cheng Tan, An-Yang Lu
arXiv:2605. 21731v2 Announce Type: replace Abstract: Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior.
By Barbara Tarantino, Gennaro Auricchio, Paolo Giudici