Learning from flowsheets: A generative transformer model for autocompletion of flowsheets
arXiv:2208. 00859v2 Announce Type: replace Abstract: We propose a novel method enabling autocompletion of chemical flowsheets.
arXiv:2208. 00778v2 Announce Type: replace-cross Abstract: SFILES are a text-based notation for chemical process flowsheets.
arXiv:2208. 00859v2 Announce Type: replace Abstract: We propose a novel method enabling autocompletion of chemical flowsheets.
arXiv:2604. 16205v2 Announce Type: replace-cross Abstract: Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, and electronic structure in chemically complex systems.
CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage.
arXiv:2608. 04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications.
arXiv:2412. 00508v2 Announce Type: replace-cross Abstract: Control structure design is an important but tedious step in P&ID development.
arXiv:2502. 18493v2 Announce Type: replace-cross Abstract: A piping and instrumentation diagram (P&ID) is a central reference document in chemical process engineering.
arXiv:2312. 02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process flows and equipment configurations.
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
arXiv:2607. 29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications.
arXiv:2608. 11283v1 Announce Type: cross Abstract: Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity.
arXiv:2607. 29479v1 Announce Type: new Abstract: Text-to-molecule generation is typically formulated as a one-shot sequence generation problem, where a model directly maps target descriptions to molecular representations.
arXiv:2605. 17758v2 Announce Type: replace Abstract: Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information.