arXiv Machine Learning By Xia Cui, Ziyi Huang, N. R. Abeynayake

Ensemble Diversity Optimization for Subjective Supervision

Read the original on arXiv Machine Learning →

arXiv:2607. 08493v1 Announce Type: new Abstract: Subjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

arXiv:2605. 02122v2 Announce Type: replace-cross Abstract: Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation.

By Akash Bonagiri, Gerard Janno Anderias, Saee Patil, Angelina Lai, Devang Borkar, Gezheng Kang, Ishant Gandhi, Setareh Rafatirad, Houman Homayoun