arXiv AI By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

Read the original on arXiv AI →

MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.

By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han