arXiv AI By Nupoor Gandhi, Emma Strubell

Task Decomposition for Efficient Annotation

Read the original on arXiv AI →

arXiv:2606. 24734v1 Announce Type: cross Abstract: High-quality annotations of structured representations are expensive to collect over large corpora.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

arXiv:2606. 02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled.

By Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
arXiv Computation and Language
Aug 25

LakeHopper: Knowledge-Aware Adaptation of Column Type Annotators across Data Lakes

LakeHopper is a method for adapting column type annotators (CTA) from one data lake to another by treating cross‑lake adaptation as a knowledge‑management problem. It decomposes the source annotator’s knowledge into source‑specific, shared, and target‑specific parts, and then uses three mechanisms—label‑set realignment, LLM‑verified gap discovery, and cluster‑based propagation with rehearsal fine‑tuning—to adapt the annotator under a limited annotation budget. The approach achieves up to a 71.4% relative macro‑F1 improvement over three PLM backbones, reaches near‑full data quality with less than 6% of target labels, and trains 27–131 times faster than fine‑tuned table LLMs.

By Yushi Sun, Xujia Li, Nan Tang, Quanqing Xu, Chuanhui Yang, Lei Chen
arXiv Computation and Language
Aug 31

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

The paper investigates whether large language models can learn and reproduce annotator‑specific label‑explanation behavior, using two sentence‑pair tasks with four annotators each. It finds that individual annotator patterns are weak at the single‑annotation level but become detectable after reducing input‑content effects and aggregating across annotators. The authors propose cross‑annotator preference optimization (CAPO), which improves upon prompting and supervised fine‑tuning by better capturing annotator‑specific reasoning while maintaining stable attribution.

By Beiduo Chen, Pingjun Hong, Ziyun Zhang, Benjamin Roth, Anna Korhonen, Barbara Plank
Hugging Face Trending Papers
Jul 30

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks.