arXiv Machine Learning

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

arXiv:2604. 26498v3 Announce Type: replace Abstract: The rapid growth of molecular foundation models and large language models (LLMs) has encouraged a scale centred view of AI in drug discovery, in which larger pretrained models are expected to supersede compact cheminformatics models.

arXiv AI
Sep 2

MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery

arXiv:2603.03517v2 Announce Type: replace-cross Abstract: General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and perfor...

By Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov, Mikolaj Mizera, Roman Schutski, Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Thomas MacDougall, Mathieu Reymond, Mihir Bafna, Kaeli Kaymak-Loveless, Eugene Babin, Maxim Malkov, Mathias Lechner, Ramin Hasani, Alexander Amini, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv Machine Learning
Jul 28

MEGA-CL: A Molecular Foundation Model for Generalizable ADMET Prediction through Graph External Attention and Contrastive Learning

arXiv:2607. 24314v1 Announce Type: new Abstract: Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery.

By Tinghui Jin, Kedu Jin, Ying Li, Guanghui Ren, Jingzhi Xue, Shiyu Zhou, Xiaoli Dai, Li-bin Wei, Xijing Chen, Di Zhao, Jinfeng Liu
arXiv AI
Aug 24

Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility

The study introduces Malaria-Instruct, a curated instruction-following dataset for malaria virtual screening, and evaluates five open-source large language models (Gemma-2, TxGemma, and LlaSMol-Mistral) against classical machine learning baselines and proprietary models. Fine‑tuned LLMs outperform all baselines, with TxGemma-9B achieving the highest ROC‑AUC (0.731 ± 0.005) and LlaSMol-Mistral-7B delivering the best enrichment factor (EF@1% ≈ 4.99). The results demonstrate that domain‑specific fine‑tuning and chemistry‑aware pretraining are essential for reliable discrimination, positioning fine‑tuned open‑source LLMs as a resource‑efficient alternative for antimalarial virtual screening.

By Marvellous O. Ajala (Magami Open Sciences Initiative), Zainab Ashimiyu-Abdusalam (Magami Open Sciences Initiative), Comfort Adesina (Magami Open Sciences Initiative)
arXiv Machine Learning
Jul 3

An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility

arXiv:2607. 02212v1 Announce Type: cross Abstract: Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both.

By Sampreeti Bhattacharya, Arkaprava Roy