arXiv AI

Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning

The paper introduces a method for estimating uncertainty in hierarchical taxonomic reasoning produced by black‑box large language models (LLMs). By extracting proxy features with an open‑source tool and training lightweight supervised estimators that incorporate hierarchy‑aware supervision, the authors predict rank‑wise correctness. Across three LLMs, these estimators outperform token‑likelihood baselines, raising micro AUROC from 0.57 to 0.75–0.80, with a rank‑specific multi‑head design delivering the best results.

arXiv Machine Learning
Sep 24

Active Learning for Biodiversity Monitoring: From Label Efficiency to Reliable Ecological Inference

The paper discusses how active learning (AL) can alleviate the expert annotation bottleneck in biodiversity monitoring by selecting the most informative samples under a fixed budget. It highlights that while AL reduces labeling effort, its non-random sample selection complicates model validation, calibration, and ecological inference, issues often overlooked in current studies. The authors review existing AL research across acoustic and image data, identify gaps such as limited species coverage and lack of real-world deployments, and propose a tutorial framework and roadmap for developing AL methods that support efficient training, reliable validation, and trustworthy ecological conclusions.

By Ben McEwen, Shiqi Zhang, Dan Stowell
arXiv Machine Learning
1d ago

Uncertainty-Aware Learning from Multi-Expert Interval Targets

The paper introduces a method for learning from multiple experts who provide interval labels, addressing both within‑label imprecision and between‑expert variation. It harmonizes diverse label vocabularies into a shared probabilistic space, retains individual intervals using a mixture of Beta distributions, and decomposes predictive uncertainty into components that are matched to their corresponding sources of label uncertainty. On sea‑ice concentration data, the approach achieves a 31% reduction in mean absolute error compared to hard‑label baselines and outperforms several aggregation and interval‑regression methods.

By Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett, Farnoush Banaei-Kashani
arXiv Machine Learning
Sep 23

Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling

The paper introduces ACORN, a method that blends machine‑learning predictions with occupancy models to guide ecologists in selecting which samples to review. By strategically choosing the most informative labels, ACORN achieves ecological conclusions nearly identical to fully human‑labeled data while dramatically reducing the number of expert reviews needed. The approach is evaluated on camera‑trap and bioacoustic datasets, demonstrating its effectiveness across real‑world biodiversity surveys.

By Timm Haucke, Lauren Harrell, Justin Kay, Mary Clapp, Sara Beery
arXiv Statistics ML
6d ago

SAGE: A sampling-aware global evaluation benchmark for species distribution modeling

The paper introduces SAGE, a Sampling‑Aware Global Evaluation benchmark for species distribution modeling that uses GBIF records for training and sPlotOpen vegetation plots for presence‑absence evaluation across 5,771 plant species. It groups species by sampling effort and relative prevalence to assess how well single‑species and multi‑species deep‑learning SDMs perform under different data conditions. The study finds that Random Forests and DeepSDMs perform best overall, with DeepSDMs excelling for infrequently recorded species only when bias‑correction techniques are applied.

By Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger, Niklaus E. Zimmermann, Lo\"ic Pellissier, Devis Tuia, Jan Dirk Wegner
arXiv Computation and Language
3d ago

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

arXiv:2607.05721v2 Announce Type: replace Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refin...

By Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu, Pei Chen, Aman Gupta, Zhe Su, Ming Tan, Zhilin Zhang, Qun Liu, Manikandarajan Ramanathan, Rajashekar Maragoud, Edward Vul, Jing Huang, Dakuo Wang
arXiv Computation and Language
Aug 28

The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions

The paper investigates the "score granularity gap" in black-box large language model (LLM) classifiers, asking how finely a confidence score can be thresholded for deployment. By comparing seven confidence construction methods across 25 model-dataset pairs, the authors find that single-shot verbalized confidence, when properly converted to a probability, ranks well but offers only a few distinct threshold values, limiting operational flexibility. The study also shows that multi-query aggregation can improve weak models but may harm strong ones, and provides concrete guidance for deployment trade-offs.

By Ao Sun, Tian Sun, Jiaxing Geng