arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song
arXiv:2511. 19264v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures.
By Amirtha Varshini A S, Duminda S. Ranasinghe, Hok Hei Tam
The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.
By Yiqi Yao, Miquel Duran-Frigola
The study fine‑tunes the Uni‑Mol2 molecular foundation model on the GS‑LF benchmark for multi‑label odor descriptor prediction. The resulting model matches or surpasses state‑of‑the‑art baselines on the primary benchmark and successfully transfers to four downstream olfactory tasks—including cross‑dataset prediction, odorless vs. odorous classification, enantiomer evaluation, and odor mixture discriminability—without further deep‑learning training. The enantiomer analysis demonstrates that 3D molecular representations can distinguish mirror‑image molecules, a capability lacking in 2D graph models, though predicting stereochemistry’s perceptual effects remains unresolved.
By Yikun Han, Yi Wang, Neil Mankodi, Stephen Yang, Ambuj Tewari
arXiv:2607. 17671v1 Announce Type: new Abstract: Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response?
By Kseniia Vaniushkina, Jeongmin Lim, Jinyong Park
HADRec is a Hierarchy-Aware Drug Recommendation framework that fuses molecular knowledge and electronic health records to improve medication recommendation. It uses LLaMA-7B to encode clinical notes, ChemBERTa to encode drug SMILES strings, and a cross‑attention mechanism for multimodal fusion, while a hierarchical predictor and consistency constraint loss enforce adherence to the ATC classification system. Experiments on MIMIC‑III and MIMIC‑IV show state‑of‑the‑art performance, strong generalization, and well‑calibrated predictions, with counterfactual evaluation indicating clinically aligned reasoning.
By Junke Wang, Hongshun Ling, Li Zhang, Jinjing Wu, Tong Shao, Fang Wang, Yuan Gao