arXiv:2609.38538v1 Announce Type: cross
Abstract: Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as...
By Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain
arXiv:2608. 06713v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected.
By Yiming Zhang, Hikaru Shindo, Shuan Chen, Kaushalya Madhawa, Jun Jin Choong, Yuna Oikawa, Takashi Fujiwara, Keisuke Ozawa
arXiv:2509. 00123v2 Announce Type: replace-cross Abstract: A fundamental challenge in microbial ecology is determining whether bacteria compete or cooperate in different environmental conditions.
By Oleksandr Cherednichenko, Josephine Solowiej-Wedderburn, Laura M. Carroll, Eric Libby
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv:2607. 20539v1 Announce Type: cross Abstract: While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited.
By Kyunghoon Hur, Eunjung Jeon, Hyun Woo Kim, Gyubok Lee, Seongjun Yang
arXiv:2609.37555v1 Announce Type: new
Abstract: Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical st...
By Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE...
arXiv:2609.06779v1 Announce Type: cross
Abstract: Drug repurposing aims to identify new therapeutic uses for existing compounds and, compared with de novo drug discovery, offers a faster and more cos...
By Zijie Liu, Hongxuan Li, Zhen Tan, Jinhao Duan, Baixiang Huang, Zunpeng Liu, Kai Shu, Tianlong Chen
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
By Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo, Tianyu Zhu, Srikanth Pilla, Gang Li, Ling Liu, Feng Luo
arXiv:2608. 02027v1 Announce Type: new Abstract: We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit.
By Jakub Adamczyk, Adam Staniszewski
arXiv:2607. 01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development.
By Wenbo Zhang
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.
By Raunak Kumar, Anuj Pal, Dhruvi Solanki, Parikshit Pareek, Juhi Singh, Jitin Singla