arXiv:2608. 19906v1 Announce Type: new Abstract: Accurately ranking active ligands for a target protein pocket from massive chemical libraries remains a central challenge in virtual screening.
By Jia-Qi Lin, Yinghua Yao, Chang-Dong Wang, Yew-Soon Ong, Yuangang Pan
TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.
By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
arXiv:2607. 17671v1 Announce Type: new Abstract: Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response?
By Kseniia Vaniushkina, Jeongmin Lim, Jinyong Park
TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.
arXiv:2510. 24380v2 Announce Type: replace Abstract: Make-on-demand combinatorial synthesis libraries (CSLs) like Enamine REAL have significantly enabled drug discovery efforts.
By Aryan Pedawi, Jordi Silvestre-Ryan, Bradley Worley, Darren J Hsu, Kushal S Shah, Elias Stehle, Jingrong Zhang, Izhar Wallach
The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.
By Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia
arXiv:2608.30877v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have demonstrated exceptional performance in protein-ligand interaction prediction, but state-of-the-ar...
By Rui Xiao, Yili Xu
arXiv:2606. 26657v1 Announce Type: new Abstract: Identifying high-utility candidates from massive discrete spaces under expensive evaluations is a recurring challenge across the sciences, with structure-based drug discovery as a prominent example.
By Mohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay, Joana Reis, Felix Strieth-Kalthoff, Haribabu Arthanari
arXiv:2606. 00555v1 Announce Type: new Abstract: Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two often-conflicting objectives -- binding affinity and druggability -- which single optimization steps rarely improve together.
By Zaifei Yang, Weiyu Chen, Yaqing Wang, James Kwok
The paper introduces a framework to predict whether a compound’s potency can be quantified in dose‑response profiling, treating quantifiability as a separate triage goal from biological activity. It shows that features from low‑cost primary screens, rather than molecular structure, strongly predict quantifiability, and that this prediction holds across new chemical scaffolds and assay families. The authors argue that incorporating quantifiability predictions can better allocate expensive dose‑response resources.
By Sean Lim
SpecOpt is a new molecular design task that optimizes the binding specificity of existing drugs by making constrained structural modifications. The method uses an agentic framework that docks a compound against its intended target and known off‑targets, compares residue‑aware atom‑protein contacts, and feeds the differential interactions to a large language model to propose changes. On a benchmark of 915 compounds, SpecOpt increased the target‑off‑target binding gap for 84.8% of cases while preserving drug‑like properties and structural similarity.
By Thao Nguyen, Heng Ji
The paper introduces GRACE, a 3D collision cross section (CCS) predictor that incorporates geometric residual adduct conditioning via early fusion. GRACE adapts a pretrained molecular geometry encoder with an adduct token and low‑rank attention adapters, achieving the lowest mean percentage differences on random, scaffold, and adduct‑sensitive splits of a curated dataset of over 9,000 experimental CCS records. Diagnostic analyses attribute its performance to residual learning that removes the dominant mass‑CCS trend and to early fusion that enhances adduct‑sensitive prediction.
By Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr., Joseph A. Morrone