arXiv:2607. 10729v1 Announce Type: new Abstract: Molecular property models are commonly evaluated by holding out Bemis--Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity.
By Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu
The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.
By Yiqi Yao, Miquel Duran-Frigola
arXiv:2607. 06605v1 Announce Type: new Abstract: Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns prediction sets containing the true label with probability at least 1 - alpha.
By Muhammadjon Tursunbadalov (School of Science and Technology, Champions College Prep, United States), Mustafojon Tursunbadalov (School of Science and Technology, Champions College Prep, United States)
arXiv:2604. 26498v3 Announce Type: replace Abstract: The rapid growth of molecular foundation models and large language models (LLMs) has encouraged a scale centred view of AI in drug discovery, in which larger pretrained models are expected to supersede compact cheminformatics models.
By Jinjiang Guo, Sheng Ding
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv:2603. 10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never tested.
By David Gringras
ToxLens is a reproducible multi‑task graph‑learning framework designed for leakage‑aware, uncertainty‑calibrated prediction of 11 molecular toxicity endpoints, including Ames mutagenicity and hERG inhibition. The workflow integrates conservative chemical curation, sphere‑exclusion filtering, a leakage‑aware UMAP‑HDBSCAN split, parallel graph and global‑feature encoders with late concatenation, temperature‑scaled Monte Carlo dropout, conformal‑style prediction sets, applicability‑domain analysis, and SHAP‑guided toxicophore discovery with occlusion controls. On a leakage‑controlled test fold, a five‑seed soft‑voting ensemble achieved MCC 0.44, AUROC 0.83, and AUPRC 0.58, outperforming four ECFP4‑based shallow baselines across all endpoints.
By Magnus H. Str{\o}mme, Alex G. C. de S\'a, David B. Ascher
arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.
By Kai Lun Huang (California State University, Fullerton), Wei Chieh Sun (University of Washington)
The study investigates whether self‑supervised pretraining improves molecular graph neural networks by adapting the LeJEPA architecture to molecular graphs. While pretraining enhances learned representations and a frozen probe outperforms random initialization on tasks such as ogbg‑molhiv, it does not consistently boost finetuning performance across different data splits. Combining pretrained embeddings with 1024‑bit Morgan fingerprints yields modest gains, indicating that pretraining provides complementary information best exploited at the feature level.
By Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki
The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.
By Bilal Ahmad, Rajed Mehmood
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
WEECFP-SuRGE introduces a 1024‑dimensional, parameter‑free continuous fingerprint that distributes each Morgan substructure across about thirty‑two signed positions in a single vector. The accompanying transformer architecture applies Substructure Rotary Graph‑distance Encoding (SuRGE), a RoPE‑like rotation based on molecular shortest‑path graph distance, to the fingerprint tokens. In benchmark tests, a seven‑model blend of this architecture achieves top rankings on the TDC ADMET leaderboard and outperforms classical fingerprints on most MoleculeNet regression tasks, while its tokenization scheme is shown to be near‑lossless and highly efficient for positional memory.
By Robert Epps