arXiv Machine Learning

Advancing Ligand-based Virtual Screening and Molecular Generation with Pretrained Molecular Embedding Distance

arXiv:2604. 24474v2 Announce Type: replace Abstract: Molecular similarity plays a central role in ligand-based drug discovery, such as virtual screening, analog searching, and goal-directed molecular generation.

arXiv Machine Learning
Jun 5

An accurate nucleic acid-small molecule docking framework via geometric deep learning with large-scale pretraining

arXiv:2606. 05198v1 Announce Type: cross Abstract: Nucleic acids are increasingly recognized as therapeutic targets beyond conventional protein-centered drug discovery, yet accurate and efficient docking of small molecules to nucleic acid structures remains challenging.

By Shi Li (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Xujun Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China), Mingquan Liu (Faculty of Health Sciences, University of Macau, Macau SAR, China), Hui Zhang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Shuoying Jia (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Yu Kang (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Shanghai Innovation Institute, Shanghai, China), Tingjun Hou (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China), Peichen Pan (College of Pharmaceutical Sciences, Zhejiang University, Hangzhou, Zhejiang, P. R. China, Zhejiang Provincial Key Laboratory for Intelligent Drug Discovery and Development, Jinhua Institute of Zhejiang University, Zhejiang, China)
arXiv AI
Aug 19

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.

By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
arXiv AI
Sep 15

Chemical and geometric representation fidelity improves drug--target affinity prediction

The paper introduces ReGeoDTA, a framework that preserves chemical heterogeneity and continuous geometric relationships in drug and protein representations to improve drug–target affinity prediction. Experiments on three benchmark datasets show that maintaining representation fidelity consistently enhances predictive accuracy across various DTA architectures, while degrading representations harms performance and cannot be recovered by more complex downstream models. The study highlights representation fidelity as a key upstream design principle for accurate and generalizable affinity prediction.

By Yixiao Li, Yining Qian, Yefan Chen, Zenghui Chen, Jiayue Sun, Yuhai Zhao, Cheng Tan, An-Yang Lu
arXiv AI
Jul 24

Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design

arXiv:2607. 20550v1 Announce Type: cross Abstract: The traditional "one drug, one target" paradigm of structure-based drug design (SBDD) frequently proves inadequate for treating multifactorial diseases such as cancer and neurodegenerative disorders, owing to compensatory signaling pathways and the emergence of drug resistance.

By Tianming Han, Zhijie Pan, Wenchi Ge, Qi Zhao
arXiv Machine Learning
5d ago

Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles

The paper presents MoCoP v2, an enhanced contrastive pretraining method that aligns small molecule embeddings with deep‑learning‑derived cell morphology profiles. By replacing CellProfiler fingerprints with richer image‑encoded features, the new embeddings better capture how molecules alter cell morphology, leading to improved QSAR, toxicity, ADME, and activity predictions. Performance scales log‑linearly with training data size, indicating further gains with larger datasets.

By Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang
arXiv AI
Sep 7

NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer

NEAT-POCKET is a pocket‑conditioned extension of the autoregressive NEAT model that generates 3D molecules atom by atom within protein binding pockets, maintaining atom permutation invariance and explicitly modeling hydrogen atoms. It outperforms existing baselines on the CrossDocked and SPINDR datasets, achieving competitive structure‑based generation performance while sampling significantly faster. The model also supports pocket‑conditioned fragment completion, a capability directly useful for lead optimization and scaffold elaboration in drug design.

By Roxane Axel Jacob, Daniel Rose, Thierry Langer, Johannes Kirchmair