The paper discusses how active learning (AL) can alleviate the expert annotation bottleneck in biodiversity monitoring by selecting the most informative samples under a fixed budget. It highlights that while AL reduces labeling effort, its non-random sample selection complicates model validation, calibration, and ecological inference, issues often overlooked in current studies. The authors review existing AL research across acoustic and image data, identify gaps such as limited species coverage and lack of real-world deployments, and propose a tutorial framework and roadmap for developing AL methods that support efficient training, reliable validation, and trustworthy ecological conclusions.
By Ben McEwen, Shiqi Zhang, Dan Stowell
arXiv:2606. 03821v1 Announce Type: new Abstract: Active learning is now standard practice in labeling ecological data, enabling ecologists to quickly process large volumes of field data to understand and monitor natural environments.
By Rupa Kurinchi-Vendhan, Sara Beery
Active learning is now standard practice in labeling ecological data, enabling ecologists to quickly process large volumes of field data to understand and monitor natural environments. Current practices evaluate active learning inductively, estimating predictive performance on a held-out test set.
arXiv:2606. 27667v1 Announce Type: cross Abstract: Artificial intelligence is transforming biodiversity monitoring by enabling automated analysis of ecological imagery collected from camera traps, drones, satellites, underwater platforms, and other sensing systems.
By Brinnae Bent, Holly R. Houliston, Jiayi Zhou, G\"unel Aghakishiyeva, David W. Johnston
arXiv:2609.15255v1 Announce Type: new
Abstract: Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However,...
By Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery
The paper introduces SAGE, a Sampling‑Aware Global Evaluation benchmark for species distribution modeling that uses GBIF records for training and sPlotOpen vegetation plots for presence‑absence evaluation across 5,771 plant species. It groups species by sampling effort and relative prevalence to assess how well single‑species and multi‑species deep‑learning SDMs perform under different data conditions. The study finds that Random Forests and DeepSDMs perform best overall, with DeepSDMs excelling for infrequently recorded species only when bias‑correction techniques are applied.
By Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger, Niklaus E. Zimmermann, Lo\"ic Pellissier, Devis Tuia, Jan Dirk Wegner