arXiv AI

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

arXiv:2605. 17254v3 Announce Type: replace Abstract: Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target properties from given structures, whereas the latter generates candidate structures according to desired properties.

arXiv Machine Learning
Sep 17

Robust and Efficient AI Frameworks for Scalable Material Design and Property Prediction

The thesis presents AI frameworks that accelerate crystalline materials discovery by tackling both crystal property prediction and crystal structure generation. It introduces CrysXPP, CrysGNN, and CrysMMNet for efficient, data‑sparse property prediction using graph autoencoding, self‑supervised pretraining, and multimodal learning. For generation, TGDMat is a text‑guided diffusion model that jointly learns lattice parameters, atomic types, and coordinates, enabling valid, stable, and conditionally generated periodic materials.

By Kishalay Das
arXiv Machine Learning
5d ago

ChemMLLM: Chemical Multimodal Large Language Model

ChemMLLM is a unified chemical multimodal large language model designed for molecule understanding and generation across text, SMILES strings, and images. The authors curated five multimodal tasks and benchmarked ChemMLLM against leading general MLLMs, chemical LLMs, and specialized models, finding it outperforms general-purpose MLLMs and matches specialized models on all tasks. The study demonstrates that a single foundation model can handle diverse cross‑modal chemical tasks, including image generation, enabling more intuitive visual human‑AI interaction.

By Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu, Shufei Zhang, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu
arXiv Machine Learning
Jun 2

Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design

arXiv:2606. 02507v1 Announce Type: cross Abstract: Inverse materials design is shifting materials discovery from forward prediction to targeted proposal of candidates that satisfy objectives under physical constraints.

By Anand Babu, Rog\'erio Almeida Gouv\^ea, Gian-Marco Rignanese
arXiv Machine Learning
Jun 17

Toward Controllable Catalyst Inverse Design via Large-Scale Autoregressive Pretraining

arXiv:2606. 17445v1 Announce Type: new Abstract: Inverse design of heterogeneous catalysts remains challenging because catalyst surfaces exhibit substantial structural complexity with coupled surface-adsorbate interactions across a vast chemical space that is difficult to explore efficiently through conventional screening alone.

By Dong Hyeon Mok, Jonggeol Na, Seoin Back
arXiv AI
Sep 23

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.

By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
arXiv AI
Aug 26

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.

By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu
arXiv Machine Learning
Jul 27

Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives

arXiv:2607. 21660v1 Announce Type: cross Abstract: Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery.

By Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, Nicola A. Morley, Robert D. J. Oliver, Alexandra J. Ramadan, Delvin Ce Zhang, Katerina A. Christofidou, Haiping Lu
arXiv Machine Learning
Sep 11

Dynamic language model representations for multi-objective reaction optimisation

The paper introduces a method that learns dynamic reaction representations directly from textual descriptions using a fine‑tuned language model coupled with Gaussian process surrogates. This approach enables multi‑objective Bayesian optimisation for chemical reactions, achieving faster convergence than traditional descriptor libraries or one‑hot encodings across nickel‑, palladium‑, and iridium‑catalysed systems. Prospective experiments on a palladium‑catalysed cyanation and an asymmetric hydrogenation produced high‑yield, high‑enantiomeric‑excess conditions after only two rounds of high‑throughput testing, translating directly to gram‑scale synthesis.

By Joshua W. Sin, David Ming Segura, Bojana Rankovi\'c, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt P\"untener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller