CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage.
arXiv:2608. 06961v1 Announce Type: new Abstract: Early-stage molecular design is an iterative process, not just a task of generating molecules.
By Zhu Wang, Jiangyu Chen, Yingjun Shang, Yuhui Yao, Laiao Lu, Tianfan Fu, Na Zou
arXiv:2607. 22677v1 Announce Type: cross Abstract: Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse.
By Sean R. Wilkinson, Polina Shpilker, Wesley Brewer
arXiv:2606. 18425v1 Announce Type: cross Abstract: Scientific workflow management systems (WMS) support scalable and reproducible execution of complex pipelines, but workflow design, implementation, and debugging remain largely manual and require significant expertise.
By Komal Thareja, Hamza Safri, Rajiv Mayani, Anirban Mandal, Ewa Deelman
arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.
By Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi, Ketan Maheshwari, Marshall McDonnell, Massimiliano Lupo Pasini, Polina Shpilker, Renan Souza, Patrick Widener, Sarp Oral, Wesley Brewer
arXiv:2608. 02027v1 Announce Type: new Abstract: We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit.
By Jakub Adamczyk, Adam Staniszewski