arXiv Computation and Language By Philipp Steigerwald, Eric Rudolph, Jens Albrecht

N\"urnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters

Read the original on arXiv Computation and Language →

The paper reports on the N"urnberg NLP team’s system for the GermEval 2026 shared task on harmful content detection in German social media. The authors tackle severe class imbalance by building a nine‑voter ensemble that varies along three orthogonal axes—LLM choice, training method, and class scope—to achieve error independence. Their system attains macro‑F1 scores of 89.56 (C2A), 71.63 (DBO), 54.84 (VIO), and 83.02 (DEF) on the hidden test set, winning all four subtasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
3d ago

StanceEval 2026: The Second Stance Detection Shared Task

StanceEval 2026 is the second edition of a shared task on stance detection in Arabic social media, where systems must classify a tweet’s stance toward a target as Favor, Against, or None. The event featured two tracks: Track 1 tests cross‑target transfer on thematically related topics (Women Driving vs. Women Empowerment), while Track 2 evaluates cross‑domain transfer to entirely unseen targets (E‑Cars and Trimester System). With 80 registered teams and 30 submissions, top systems achieved $F_{avg2}$ scores of 0.8994 (Track 1) and 0.9400 (Track 2), surpassing baseline performance and highlighting challenges such as target polarization and dialectal nuance.

By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif
Hugging Face Trending Papers
Aug 11

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6.

arXiv AI
Sep 24

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

arXiv:2609.27165v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix backgr...

By Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang