arXiv Machine Learning

ALF: An Active Learning Framework for Scientific Discovery

ALF is a modular active learning framework designed to streamline the entire data acquisition process for scientific discovery. It offers a single API that supports both offline benchmarking against existing datasets and online deployment with an oracle for real‑world candidate acquisition. The framework is open‑source and available on GitHub.

arXiv AI
Jul 7

Automated Data Readiness for Scientific AI

arXiv:2607. 02771v1 Announce Type: new Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data.

By Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi, Ketan Maheshwari, Marshall McDonnell, Massimiliano Lupo Pasini, Polina Shpilker, Renan Souza, Patrick Widener, Sarp Oral, Wesley Brewer
arXiv Machine Learning
Sep 15

BioDCASE: Active Learning for Bioacoustics

arXiv:2609.15255v1 Announce Type: new Abstract: Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However,...

By Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery
arXiv Machine Learning
Aug 27

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

Flower Hub is a platform that allows researchers to publish, discover, and run federated learning (FL) benchmarks in a reproducible way. It packages benchmarks as executable, versioned applications with standardized metadata, pinned dependencies, and explicit evaluation workflows, enabling the same benchmark to run in both simulation and real deployment environments. The platform includes a multi-domain benchmark suite covering cross-silo and cross-device settings in areas such as medical imaging, finance, legal instruction tuning, phishing detection, and audio tagging, and it supports system-aware reporting of runtime and communication metrics.

By Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane
arXiv Machine Learning
Jun 25

Why Pool When You Can Flow? Active Learning with GFlowNets

arXiv:2509. 00704v2 Announce Type: replace Abstract: The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screening for drug discovery.

By Renfei Zhang, Mohit Pandey, Artem Cherkasov, Martin Ester
arXiv AI
Jun 2

CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing

arXiv:2605. 11359v3 Announce Type: replace Abstract: Scientific data processing often requires task-specific algorithms or AI models, creating a barrier for domain scientists who need to analyze their data but may not have extensive computing or image-processing expertise.

By Ming Du, Xiangyu Yin, Yanqi Luo, Dishant Beniwal, Songyuan Tang, Hemant Sharma, Mathew J. Cherukara
arXiv AI
Jun 4

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.

By Trinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, Thanh Tam Nguyen