arXiv AI By Konstantinos Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras

Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools

Read the original on arXiv AI →

arXiv:2607. 13115v1 Announce Type: new Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

WOMBAT: Whitebox Oracle for Molecular Benchmarking and Attribution Testing

WOMBAT is a benchmark comprising 14 whitebox graph neural networks (GNNs) whose message‑passing weights are manually set to detect specific SMARTS motifs. Each model’s decision rule is explicitly known, providing a ground truth for attribution that allows researchers to identify and study errors in post‑hoc explainers such as GNNExplainer, PGExplainer, and Integrated Gradients. The authors validate the models on millions of PubChem molecules, demonstrate how Integrated Gradients can be misled to spread attribution, and release the dataset, models, and evaluation code for future XAI tool development.

By Dominik Matuszek, Bartosz Zieli\'nski, Tomasz Danel, Dawid Rymarczyk
arXiv Machine Learning
Jun 5

MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.

By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han