arXiv:2606. 25989v1 Announce Type: cross Abstract: Automated classification of marine species from underwater imagery is essential for scalable ocean biodiversity monitoring and conservation policy.
By Dan Zimmerman, Dimitris A. Pados, George Sklivanitis
The paper discusses how active learning (AL) can alleviate the expert annotation bottleneck in biodiversity monitoring by selecting the most informative samples under a fixed budget. It highlights that while AL reduces labeling effort, its non-random sample selection complicates model validation, calibration, and ecological inference, issues often overlooked in current studies. The authors review existing AL research across acoustic and image data, identify gaps such as limited species coverage and lack of real-world deployments, and propose a tutorial framework and roadmap for developing AL methods that support efficient training, reliable validation, and trustworthy ecological conclusions.
By Ben McEwen, Shiqi Zhang, Dan Stowell
The paper introduces a method for learning from multiple experts who provide interval labels, addressing both within‑label imprecision and between‑expert variation. It harmonizes diverse label vocabularies into a shared probabilistic space, retains individual intervals using a mixture of Beta distributions, and decomposes predictive uncertainty into components that are matched to their corresponding sources of label uncertainty. On sea‑ice concentration data, the approach achieves a 31% reduction in mean absolute error compared to hard‑label baselines and outperforms several aggregation and interval‑regression methods.
By Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett, Farnoush Banaei-Kashani
The paper introduces ACORN, a method that blends machine‑learning predictions with occupancy models to guide ecologists in selecting which samples to review. By strategically choosing the most informative labels, ACORN achieves ecological conclusions nearly identical to fully human‑labeled data while dramatically reducing the number of expert reviews needed. The approach is evaluated on camera‑trap and bioacoustic datasets, demonstrating its effectiveness across real‑world biodiversity surveys.
By Timm Haucke, Lauren Harrell, Justin Kay, Mary Clapp, Sara Beery
arXiv:2609.11986v1 Announce Type: cross
Abstract: Passive acoustic monitoring produces far more bat recordings than experts can label. We show that simple model-generated pseudo-labels turn this surp...
By Frank Fundel, Alexandra Howard
arXiv:2608. 08398v1 Announce Type: new Abstract: Astronomers classify galaxy morphology to investigate cosmic evolution.
By Kai Cheng, Ruoqi Wang, Qiong Luo
The paper introduces SAGE, a Sampling‑Aware Global Evaluation benchmark for species distribution modeling that uses GBIF records for training and sPlotOpen vegetation plots for presence‑absence evaluation across 5,771 plant species. It groups species by sampling effort and relative prevalence to assess how well single‑species and multi‑species deep‑learning SDMs perform under different data conditions. The study finds that Random Forests and DeepSDMs perform best overall, with DeepSDMs excelling for infrequently recorded species only when bias‑correction techniques are applied.
By Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger, Niklaus E. Zimmermann, Lo\"ic Pellissier, Devis Tuia, Jan Dirk Wegner
arXiv:2608. 06406v1 Announce Type: cross Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management.
By Laura Bader, Muhammad Ammar Ahmed, Xiao Xiang Zhu, G\"oran Kauermann
Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors.
arXiv:2607.05721v2 Announce Type: replace
Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refin...
By Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu, Pei Chen, Aman Gupta, Zhe Su, Ming Tan, Zhilin Zhang, Qun Liu, Manikandarajan Ramanathan, Rajashekar Maragoud, Edward Vul, Jing Huang, Dakuo Wang
The paper investigates the "score granularity gap" in black-box large language model (LLM) classifiers, asking how finely a confidence score can be thresholded for deployment. By comparing seven confidence construction methods across 25 model-dataset pairs, the authors find that single-shot verbalized confidence, when properly converted to a probability, ranks well but offers only a few distinct threshold values, limiting operational flexibility. The study also shows that multi-query aggregation can improve weak models but may harm strong ones, and provides concrete guidance for deployment trade-offs.
By Ao Sun, Tian Sun, Jiaxing Geng
arXiv:2607. 02182v1 Announce Type: new Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence, severely hindering trustworthy deployment.
By Jijie Zhang, Zhe Ren, Quan Zhang, Dandan Guo