arXiv:2604.13899v5 Announce Type: replace-cross
Abstract: Annotating data remains a costly bottleneck for supervised NLP. Active learning (AL) reduces the number of human labels needed by selecting o...
By Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein, Hinrich Sch\"utze
arXiv:2606. 08718v1 Announce Type: cross Abstract: While Deep Active Learning (DAL) effectively reduces human annotation costs, its efficacy is constrained by human annotation errors.
By Md Abdullah Al Forhad, Weishi Shi
arXiv:2601.14172v4 Announce Type: replace-cross
Abstract: We study neural multi-label classification under severe label imbalance through sentence-level detection of the 19 refined Schwartz human val...
By V\'ictor Yeste, Paolo Rosso
arXiv:2608. 09209v1 Announce Type: cross Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs.
By Chidaksh Ravuru, Shashank Srivastava
arXiv:2609.22133v1 Announce Type: new
Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at...
By Kentaro Nakamura, Jing Ling Tan, George Yean
arXiv:2609.38630v1 Announce Type: new
Abstract: Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with...
By Jonathan Graehl