arXiv Machine Learning By Tianren Zhang

Molecular representation shapes the balance between target fidelity and exploration in flow based polymer generation

Read the original on arXiv Machine Learning →

The paper introduces PolyLatentFlow, a continuous‑time flow‑matching framework for polymer generation, and LlamaUni, a multimodal representation that fuses polymer sequences with 3D structural data. In unconditional generation, the combination yields the highest number of valid, novel candidates while preserving diversity, and in conditional settings it systematically shifts property distributions across a 200 °C target range. Across multi‑property tasks, the representation choice affects validity, training‑set replay, and structural proximity, with PolyLatentFlow + LlamaUni achieving the best balance of high validity, low replay, and high target hit yield.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design

HiPoly is a polymer-native AI framework that uses a three-level hierarchical graph architecture built on the G2RINS representation to process complete polymer descriptions. It encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, enabling end-to-end workflows from experimental data to property prediction, generative design, and physics-based validation. The framework achieves state-of-the-art accuracy for thermophysical properties of multi-component polymer systems and demonstrates generative design by discovering sustainable, PFAS-free alternatives with target surface-energy properties.

By Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Dasha Safarian, Ming Han, Juan J. de Pablo
Hugging Face Trending Papers
Sep 8

Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules

The paper introduces Equivariant-Free Transformer-Autoencoded Latent Flow Matching (EF‑TALFM), a two‑stage generative framework that uses a single fixed‑dimensional latent vector to produce variable‑size 3D molecules. The first stage samples the latent vector via flow matching, and the second stage employs an autoregressive Transformer decoder that determines molecule size while generating atom types, coordinates, and chemical states. EF‑TALFM outperforms prior methods on the PCQM4Mv2 benchmark, achieving higher uniqueness, novelty, and computational throughput, and its internal ranking improves the hit rate for target HOMO–LUMO gaps while maintaining novelty.