arXiv Machine Learning

Graph Set Transformer

arXiv:2606. 05116v1 Announce Type: new Abstract: We introduce the Graph Set Transformer (GST), a neural network architecture for learning on sets of graphs, designed for tasks in which per-element predictions depend on set-wide context as well as local structure.

arXiv Machine Learning
Jul 22

GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks

arXiv:2607. 19083v1 Announce Type: new Abstract: Equivariant graph neural networks provide a powerful modeling language for three-dimensional scientific data, but their reuse is often limited by implementations tied to specific tasks, outputs, and training regimes.

By Daniele Angioletti, Marco Nobile, Vittorio Limongelli
arXiv Machine Learning
Sep 22

Role-Aware Morgan Fingerprints for Reaction Yield Prediction

The paper introduces MFP, a reaction yield prediction method that uses role-aware Morgan fingerprints. It computes count-based circular fingerprints for each reaction component, aggregates them by chemical role, and combines them with transformation-sensitive difference features into a fixed-length descriptor for a feed-forward neural regressor. On the Suzuki‑Miyaura and Buchwald‑Hartwig benchmarks, MFP achieves R² scores of 0.878 and 0.969 respectively, while training an order of magnitude faster than graph or Transformer-based alternatives.

By Chinmay Mirji, Prashant Shekhar, Foram Madiyar, Hao Peng
arXiv Machine Learning
Sep 17

ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs

The paper introduces ADAPT, a lightweight machine‑learning force field that replaces graph neural networks with a direct coordinates‑in‑space Transformer encoder to model all pairwise atomic interactions. Applied to silicon point defects, ADAPT reduces force prediction error by about 22% and energy prediction error by roughly 40% compared to a state‑of‑the‑art GNN model, while also cutting computational cost. This approach addresses common GNN issues such as oversmoothing, oversquashing, and poor long‑range interaction representation, which are especially problematic for point defect modeling.

By Evan Dramko, Yihuang Xiong, Yizhi Zhu, Geoffroy Hautier, Thomas Reps, Christopher Jermaine, Anastasios Kyrillidis
arXiv Machine Learning
Aug 12

Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching

arXiv:2602. 13136v2 Announce Type: replace Abstract: Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization.

By Chenguang Wang, Zihan Zhou, Lei Bai, Tianshu Yu
arXiv Machine Learning
Sep 17

Robust and Efficient AI Frameworks for Scalable Material Design and Property Prediction

The thesis presents AI frameworks that accelerate crystalline materials discovery by tackling both crystal property prediction and crystal structure generation. It introduces CrysXPP, CrysGNN, and CrysMMNet for efficient, data‑sparse property prediction using graph autoencoding, self‑supervised pretraining, and multimodal learning. For generation, TGDMat is a text‑guided diffusion model that jointly learns lattice parameters, atomic types, and coordinates, enabling valid, stable, and conditionally generated periodic materials.

By Kishalay Das