arXiv Machine Learning

OPDiv: Optimal Selection of Top-K High-Scoring, Diverse Compounds

OPDiv is a new algorithm that addresses the tradeoff between ranking quality and chemical diversity in virtual screening. It uses integer optimization to select an optimal subset of top‑k compounds that meet a specified diversity threshold, evaluated with fingerprint, shape, and electrostatic metrics. The method demonstrates how to benchmark virtual screening pipelines by comparing their best achievable diverse selections.

arXiv Machine Learning
Jun 26

Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical Space

arXiv:2606. 26657v1 Announce Type: new Abstract: Identifying high-utility candidates from massive discrete spaces under expensive evaluations is a recurring challenge across the sciences, with structure-based drug discovery as a prominent example.

By Mohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay, Joana Reis, Felix Strieth-Kalthoff, Haribabu Arthanari
arXiv Machine Learning
Sep 7

Small Molecule Optimization with Large Language Models

The paper introduces Mol-E, an evolutionary algorithm that leverages large language models trained on molecular data to generate candidate molecules. Mol-E achieves state‑of‑the‑art performance on the Practical Molecular Optimization benchmark, scoring 17.500 in the task‑agnostic regime and 20.551 in the task‑informed regime. It also outperforms baseline methods in multi‑property optimization tasks involving docking against DRD2, MK2, and AChE.

By Philipp Guevorguian, Menua Bedrosian, Tigran Fahradyan, Gayane Chilingaryan, Armen Aghajanyan, Hrant Khachatrian
arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

arXiv Machine Learning
Jul 16

HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

arXiv:2607. 13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo.

By Daria A. Ryabchenko (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Pavel Gurevich (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Shamil Kadyrov (Ligand Pro, Moscow, Russia), Daria Frolova (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Kseniia Fedisheva (Ligand Pro, Moscow, Russia), Sergei A. Nikolenko (Ligand Pro, Moscow, Russia), Alexander Shapeev (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Marina A. Pak (Ligand Pro, Moscow, Russia)