SpecOpt is a new molecular design task that optimizes the binding specificity of existing drugs by making constrained structural modifications. The method uses an agentic framework that docks a compound against its intended target and known off‑targets, compares residue‑aware atom‑protein contacts, and feeds the differential interactions to a large language model to propose changes. On a benchmark of 915 compounds, SpecOpt increased the target‑off‑target binding gap for 84.8% of cases while preserving drug‑like properties and structural similarity.
By Thao Nguyen, Heng Ji
ProbeMatchDTI is a new framework for drug‑target interaction prediction that uses probe‑driven pattern matching to preserve weak biochemical signals. It introduces IterProbe, which retains contextual states across refinement depths and selects them with learnable probes, and BindingProbe, which models drug‑protein complementarity at both local and whole‑pair levels. Experiments show that ProbeMatchDTI outperforms existing methods, improving AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank, and its predictions can be integrated into downstream drug‑discovery workflows.
By Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
arXiv:2507.07426v4 Announce Type: replace
Abstract: Recent advances in large language models have demonstrated considerable potential in scientific domains such as drug repositioning. However, their...
By Zerui Yang, Yuwei Wan, Siyu Yan, Yudai Matsuda, Tong Xie, Linqi Song
arXiv:2607. 19044v1 Announce Type: new Abstract: Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design.
By Mingxuan Ouyang, Hao Lan, Wanyu Lin
ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv:2606. 01461v1 Announce Type: new Abstract: Developing effective anticancer therapeutics remains challenging due to tumor heterogeneity and the absence of well-defined molecular targets across cancer subtypes.
By Brenda Nogueira, Gisela A. Gonzalez-Montiel, Nitesh V. Chawla, Nuno Moniz
arXiv:2511.20510v3 Announce Type: replace
Abstract: Molecule generation from extremely limited training data is a key challenge in drug discovery. Existing fragment-based methods are more suitable th...
By Yuto Suzuki, Paul Awolade, Daniel V. LaBarbera, Farnoush Banaei-Kashani
arXiv:2609.00189v1 Announce Type: new
Abstract: Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implement...
By Shiyun Wa, Yifei Wang, Anna G. Green, Simone Sciabola, Ye Wang
VINCENT is a post‑training framework that provides validated, chemically coherent explanations for drug synergy predictions by extracting atom‑pair evidence from attention and gradient signals, grouping them into motifs, and refining these motifs through repeated local perturbations. On a literature‑annotated subset of 25 drug pairs, VINCENT achieves a mean motif recall of 0.826, outperforming baselines (0.49–0.66). Across 71 test pairs, its validated interaction scores yield a TP/TN separation of 3.36, indicating more accurate recovery of literature‑supported molecular regions and better alignment with predictor behavior.
By Fan-Sheng Chuang, Xuchen Li, Yujing Bian, Kaixiong Zhou
arXiv:2606. 01220v1 Announce Type: cross Abstract: Generating molecules that simultaneously satisfy drug-like properties and conform to the 3D structure of a target protein is a core challenge in structure-based drug design (SBDD).
By Guang Lin, Shikui Tu, Lei Xu
arXiv:2607. 12349v1 Announce Type: new Abstract: Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design.
By Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen, Frazier N. Baker, David C. Kombo, John L. Kane Jr., Andrew A. Scholte, Yi Li, Matthew J. LaMarche, Luigi I. Iconaru, Hans-Peter Biemann, Mingyi Hong, Xia Ning