arXiv:2603.03517v2 Announce Type: replace-cross
Abstract: General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and perfor...
By Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov, Mikolaj Mizera, Roman Schutski, Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Thomas MacDougall, Mathieu Reymond, Mihir Bafna, Kaeli Kaymak-Loveless, Eugene Babin, Maxim Malkov, Mathias Lechner, Ramin Hasani, Alexander Amini, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv:2608. 11444v1 Announce Type: cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs.
By Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv:2609.37555v1 Announce Type: new
Abstract: Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical st...
By Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv:2607. 24314v1 Announce Type: new Abstract: Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery.
By Tinghui Jin, Kedu Jin, Ying Li, Guanghui Ren, Jingzhi Xue, Shiyu Zhou, Xiaoli Dai, Li-bin Wei, Xijing Chen, Di Zhao, Jinfeng Liu
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
The study introduces Malaria-Instruct, a curated instruction-following dataset for malaria virtual screening, and evaluates five open-source large language models (Gemma-2, TxGemma, and LlaSMol-Mistral) against classical machine learning baselines and proprietary models. Fine‑tuned LLMs outperform all baselines, with TxGemma-9B achieving the highest ROC‑AUC (0.731 ± 0.005) and LlaSMol-Mistral-7B delivering the best enrichment factor (EF@1% ≈ 4.99). The results demonstrate that domain‑specific fine‑tuning and chemistry‑aware pretraining are essential for reliable discrimination, positioning fine‑tuned open‑source LLMs as a resource‑efficient alternative for antimalarial virtual screening.
By Marvellous O. Ajala (Magami Open Sciences Initiative), Zainab Ashimiyu-Abdusalam (Magami Open Sciences Initiative), Comfort Adesina (Magami Open Sciences Initiative)
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
arXiv:2609.38744v1 Announce Type: new
Abstract: Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials de...
By Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo
arXiv:2606. 19245v1 Announce Type: new Abstract: Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions.
By Hannah Le, Ramesh Ramasamy, Alex Urrutia, Mahsa Yazdani, Tim Proctor, Kenny Workman
arXiv:2607. 02212v1 Announce Type: cross Abstract: Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both.
By Sampreeti Bhattacharya, Arkaprava Roy