HiPoly is a polymer-native AI framework that uses a three-level hierarchical graph architecture built on the G2RINS representation to process complete polymer descriptions. It encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, enabling end-to-end workflows from experimental data to property prediction, generative design, and physics-based validation. The framework achieves state-of-the-art accuracy for thermophysical properties of multi-component polymer systems and demonstrates generative design by discovering sustainable, PFAS-free alternatives with target surface-energy properties.
By Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Dasha Safarian, Ming Han, Juan J. de Pablo
The paper introduces PolyLatentFlow, a continuous‑time flow‑matching framework for polymer generation, and LlamaUni, a multimodal representation that fuses polymer sequences with 3D structural data. In unconditional generation, the combination yields the highest number of valid, novel candidates while preserving diversity, and in conditional settings it systematically shifts property distributions across a 200 °C target range. Across multi‑property tasks, the representation choice affects validity, training‑set replay, and structural proximity, with PolyLatentFlow + LlamaUni achieving the best balance of high validity, low replay, and high target hit yield.
By Tianren Zhang
arXiv:2605. 26833v2 Announce Type: replace-cross Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging.
By Yasharth Yadav, Tze Kwang Gerald Er, Atsushi Goto, Kelin Xia
arXiv:2606. 20753v2 Announce Type: replace-cross Abstract: Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing.
By Chenyao Ma, Linda Zhang, Yuheng Chen, Wei Du, Shangwen Fang, Zihao Jiang, Chuanyu Liu, Xinyu Ma, Rui Su, Gang Wang, Muyao Yu, Dong Zhong, Jie Zhu, Weibo Gong, Huan Gu, Limin Li, Chen Shen, Rui Wu, Zhenghao Wu, Kan Xu, Min Zhou, Donglin He, Xiayun Huang, Shan Jiang, Pengfei Ou, Jiayu Peng, Yuwei Zhang, Jie Zhao, Di Zhang, Piao Ma, Zhenghao Li, Hao Li
The paper proposes a new framework called ensemble-conditioned guidance that reframes molecular design as an optimisation over both the modes and properties of a molecule’s conformational ensemble. It allows 3D generative models to be conditioned simultaneously on multiple axes—such as shapes, pharmacophore profiles, or protein pockets—by adaptively combining vector fields from each condition. The authors introduce adaptive symmetry learning for composable conditions across reference frames, extend the framework to support flexible-size generation, and demonstrate its effectiveness on new benchmarks and practical drug‑discovery tasks, showing improved outcomes when conditioning on additional states compared to single‑state approaches.
By Ross Irwin, Alessandro Tibo, Jon Paul Janet, Simon Olsson
arXiv:2608. 01431v1 Announce Type: cross Abstract: Polymer property prediction and inverse generative design targeting desired properties are two crucial tasks in machine learning-assisted polymer design.
By Charlie Pyle, Adarsh Gadari, C. Adrian Figg, Zhenquan Jia, Yaohang Li, Chunjiang Zhu
arXiv:2508. 02641v2 Announce Type: replace-cross Abstract: Molecular crystal structure prediction (CSP) is essential for applications in pharmaceuticals and organic electronics.
By Vahe Gharakhanyan, Yi Yang, Luis Barroso-Luque, Daniel S. Levine, Sushree Jagriti Sahoo, Brandon M. Wood, Kyle Michel, Muhammed Shuaibi, Gregory J. O. Beran, Viachaslau Bernat, Misko Dzamba, Xiang Fu, Meng Gao, Xingyu Liu, Benjamin K. Miller, Keian Noori, Lafe J. Purvis, Tingling Rao, Ammar Rizvi, Matt Uyttendaele, Andrew J. Ouderkirk, Chiara Daraio, C. Lawrence Zitnick, Arman Boromand, Noa Marom, Zachary W. Ulissi, Anuroop Sriram
Polymer property prediction and inverse generative design targeting desired properties are two crucial tasks in machine learning-assisted polymer design. While the former has received considerable attention, there have been limited methods developed for the latter.
arXiv:2609.39773v1 Announce Type: new
Abstract: Crystal structure governs material properties, making crystal structure prediction (CSP) a fundamental problem in materials science. Generative models...
By Thomas Egg, Harry Winston Sullivan, Maya M. Martirossyan, Philipp H\"ollmer, Cheng Zeng, Adrian Roitberg, Mingjie Liu, Richard Hennig, Sapna Sarupria, Ellad B. Tadmor, Stefano Martiniani
The paper introduces a chemical language foundation model that uses a SMILES‑based polymer graph representation (CPG) to encode polymer architecture and connectivity, addressing gaps in existing line notations. The model achieves strong performance across 30 polymer property benchmarks and demonstrates robustness to structural representation perturbations, with even chemically invalid SMILES sometimes matching state‑of‑the‑art results. Control experiments and attention analyses confirm that CPG offers meaningful advantages while highlighting the model’s ability to interpolate SMILES sequence space in a way that aligns loosely with chemical and architectural space.
By Nathaniel H. Park, Eduardo Soares, Victor Y. Shirasuna, Tiffany J. Callahan, Sara Capponi, Emilio Vital Brazil
arXiv:2410.20317v2 Announce Type: replace
Abstract: Molecular dynamics (MD) simulations are a principled but computationally expensive approach for studying protein conformational variability, making...
By Siddharth Viswanath, Xingzhi Sun, Lucas Lee, Danqi Liao, Hiren Madhu, David R. Johnson, Jo\~ao Felipe Rocha, Egbert Castro, Jackson D. Grady, Michael Perlmutter, Dhananjay Bhaskar, Smita Krishnaswamy
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar