GraphNOSE: A Graph Transformer in Olfaction
arXiv:2609.05694v1 Announce Type: new Abstract: Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to o...
The study fine‑tunes the Uni‑Mol2 molecular foundation model on the GS‑LF benchmark for multi‑label odor descriptor prediction. The resulting model matches or surpasses state‑of‑the‑art baselines on the primary benchmark and successfully transfers to four downstream olfactory tasks—including cross‑dataset prediction, odorless vs. odorous classification, enantiomer evaluation, and odor mixture discriminability—without further deep‑learning training. The enantiomer analysis demonstrates that 3D molecular representations can distinguish mirror‑image molecules, a capability lacking in 2D graph models, though predicting stereochemistry’s perceptual effects remains unresolved.
arXiv:2609.05694v1 Announce Type: new Abstract: Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to o...
arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.
The paper introduces a bio‑inspired deep learning framework that models olfactory perception of complex chemical mixtures. It constructs neural response curves for molecule‑receptor interactions, fuses attention‑weighted multi‑receptor and concentration‑dependent multi‑molecule data, and transfers knowledge from molecular associations to improve mixture recognition. The model achieves 92.2% accuracy and offers a generalizable computational pathway from chemical blending to perceptual formation.
arXiv:2606. 11508v1 Announce Type: new Abstract: Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy, interdependent, and often data-limited.
arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
arXiv:2510. 14217v2 Announce Type: replace Abstract: The spectral properties of feature embeddings offer critical insights into model generalization and representation quality.
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
arXiv:2510.07289v2 Announce Type: replace Abstract: Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated s...
arXiv:2609.15611v1 Announce Type: cross Abstract: Molecular property prediction requires representations that generalize from limited labeled data to structurally novel compounds. Existing molecular...
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.