arXiv AI

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.

arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv Machine Learning
Aug 27

A General-Purpose Molecular Foundation Model Transfers Across Diverse Olfactory Tasks

The study fine‑tunes the Uni‑Mol2 molecular foundation model on the GS‑LF benchmark for multi‑label odor descriptor prediction. The resulting model matches or surpasses state‑of‑the‑art baselines on the primary benchmark and successfully transfers to four downstream olfactory tasks—including cross‑dataset prediction, odorless vs. odorous classification, enantiomer evaluation, and odor mixture discriminability—without further deep‑learning training. The enantiomer analysis demonstrates that 3D molecular representations can distinguish mirror‑image molecules, a capability lacking in 2D graph models, though predicting stereochemistry’s perceptual effects remains unresolved.

By Yikun Han, Yi Wang, Neil Mankodi, Stephen Yang, Ambuj Tewari
arXiv Machine Learning
1d ago

HADRec: A Hierarchy-Aware Drug Recommendation Framework by Fusing Molecular Knowledge and Electronic Health Record

HADRec is a Hierarchy-Aware Drug Recommendation framework that fuses molecular knowledge and electronic health records to improve medication recommendation. It uses LLaMA-7B to encode clinical notes, ChemBERTa to encode drug SMILES strings, and a cross‑attention mechanism for multimodal fusion, while a hierarchical predictor and consistency constraint loss enforce adherence to the ATC classification system. Experiments on MIMIC‑III and MIMIC‑IV show state‑of‑the‑art performance, strong generalization, and well‑calibrated predictions, with counterfactual evaluation indicating clinically aligned reasoning.

By Junke Wang, Hongshun Ling, Li Zhang, Jinjing Wu, Tong Shao, Fang Wang, Yuan Gao
arXiv Machine Learning
Sep 14

Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion

The paper introduces GRACE, a 3D collision cross section (CCS) predictor that incorporates geometric residual adduct conditioning via early fusion. GRACE adapts a pretrained molecular geometry encoder with an adduct token and low‑rank attention adapters, achieving the lowest mean percentage differences on random, scaffold, and adduct‑sensitive splits of a curated dataset of over 9,000 experimental CCS records. Diagnostic analyses attribute its performance to residual learning that removes the dominant mass‑CCS trend and to early fusion that enhances adduct‑sensitive prediction.

By Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr., Joseph A. Morrone
arXiv Machine Learning
Sep 22

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.

By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
Hugging Face Trending Papers
Jul 6

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles. Preclinical transfer-learning models can simulate drug-induced expression changes but are often hard to interpret and unstable, whereas knowledge-graph methods provide mechanistic context yet remain static and fail to capture drug-induced transcriptomic perturbation dynamics.