Automated Data Readiness for Scientific AI
arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.
ALF is a modular active learning framework designed to streamline the entire data acquisition process for scientific discovery. It offers a single API that supports both offline benchmarking against existing datasets and online deployment with an oracle for real‑world candidate acquisition. The framework is open‑source and available on GitHub.
arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.
arXiv:2609.15255v1 Announce Type: new Abstract: Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However,...
Flower Hub is a platform that allows researchers to publish, discover, and run federated learning (FL) benchmarks in a reproducible way. It packages benchmarks as executable, versioned applications with standardized metadata, pinned dependencies, and explicit evaluation workflows, enabling the same benchmark to run in both simulation and real deployment environments. The platform includes a multi-domain benchmark suite covering cross-silo and cross-device settings in areas such as medical imaging, finance, legal instruction tuning, phishing detection, and audio tagging, and it supports system-aware reporting of runtime and communication metrics.
arXiv:2607. 22677v1 Announce Type: cross Abstract: Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse.
arXiv:2607. 28990v1 Announce Type: new Abstract: Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim.
arXiv:2608. 04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications.
arXiv:2608. 03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required.
arXiv:2512.12870v2 Announce Type: replace-cross Abstract: Active Learning (AL) is commonly used in applications where labeling data is expensive or time-consuming. In practice, however, labels are of...
arXiv:2509. 00704v2 Announce Type: replace Abstract: The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screening for drug discovery.
arXiv:2605. 11359v3 Announce Type: replace Abstract: Scientific data processing often requires task-specific algorithms or AI models, creating a barrier for domain scientists who need to analyze their data but may not have extensive computing or image-processing expertise.
arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.
arXiv:2603. 17216v2 Announce Type: replace Abstract: With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal.