arXiv Computation and Language

N\"urnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters

The paper reports on the N"urnberg NLP team’s system for the GermEval 2026 shared task on harmful content detection in German social media. The authors tackle severe class imbalance by building a nine‑voter ensemble that varies along three orthogonal axes—LLM choice, training method, and class scope—to achieve error independence. Their system attains macro‑F1 scores of 89.56 (C2A), 71.63 (DBO), 54.84 (VIO), and 83.02 (DEF) on the hidden test set, winning all four subtasks.

arXiv Computation and Language
3d ago

StanceEval 2026: The Second Stance Detection Shared Task

StanceEval 2026 is the second edition of a shared task on stance detection in Arabic social media, where systems must classify a tweet’s stance toward a target as Favor, Against, or None. The event featured two tracks: Track 1 tests cross‑target transfer on thematically related topics (Women Driving vs. Women Empowerment), while Track 2 evaluates cross‑domain transfer to entirely unseen targets (E‑Cars and Trimester System). With 80 registered teams and 30 submissions, top systems achieved $F_{avg2}$ scores of 0.8994 (Track 1) and 0.9400 (Track 2), surpassing baseline performance and highlighting challenges such as target polarization and dialectal nuance.

By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif
Hugging Face Trending Papers
Aug 11

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6.

arXiv AI
Sep 24

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

arXiv:2609.27165v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix backgr...

By Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang
arXiv Machine Learning
5d ago

Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

The paper presents a system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media. It tackles three tasks—risk-level classification, evidence phrase extraction, and multi-label factor identification—using Qwen2.5-Instruct models adapted with quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. The final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2, demonstrating that task‑specific training and tailored aggregation improve performance across the three tasks.

By Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
arXiv AI
Sep 25

An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection

The paper presents an explainable hate‑speech detection framework that combines DistilBERT embeddings, a Bi‑LSTM network, and an attention mechanism to capture contextual and sequential information. It uses LIME to highlight influential text features, providing transparency in predictions. Evaluated on two benchmark datasets for both binary and multi‑class tasks, the model achieves F1‑scores of 96.78%–99.53% for binary classification and 94.99%–97.00% for multi‑class classification, outperforming existing baselines.

By Rameesha Zia, Muhammad Shahid Iqbal Malik