arXiv Computation and Language

Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety

arXiv Computation and Language
2d ago

Automatic Evaluation of Mental Health Stigma in Online Communication

The paper presents a new benchmark for automatically evaluating mental health stigma in online text, featuring a fine‑grained taxonomy that covers stigma mode, domain, and specific components across multiple mental health conditions. The authors annotate naturally occurring news and social media posts and test large language models and classifiers for sentiment, toxicity, and hate speech, finding that these models poorly capture stigma and often overpredict it without explicit rules. The benchmark, annotations, exemplar cases, and code are publicly released on GitHub.

By Naomi Baes, Jemima Kang, Nick Haslam, Chris Groot, Alsa Wu, Luc Raszewski, Yulia Otmakhova
arXiv Computation and Language
Aug 31

Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech

The paper introduces FAID, a fine‑grained adaptive framework for detecting implicit hate speech. It first classifies samples into Shallow, Targeted, or Context‑Dependent categories and then applies tailored strategies—prompt‑tuning for shallow cases, knowledge augmentation for targeted ones, and an agentic prompt‑generation system for context‑dependent posts. Experiments on four benchmark datasets show that FAID outperforms state‑of‑the‑art baselines by allocating computational effort only where needed.

By Han Wang, Yuhu Cheng, Xuesong Wang, Yi Zhu
arXiv AI
Sep 25

An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection

The paper presents an explainable hate‑speech detection framework that combines DistilBERT embeddings, a Bi‑LSTM network, and an attention mechanism to capture contextual and sequential information. It uses LIME to highlight influential text features, providing transparency in predictions. Evaluated on two benchmark datasets for both binary and multi‑class tasks, the model achieves F1‑scores of 96.78%–99.53% for binary classification and 94.99%–97.00% for multi‑class classification, outperforming existing baselines.

By Rameesha Zia, Muhammad Shahid Iqbal Malik
arXiv Computation and Language
Sep 18

Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

The paper introduces ProKDA, a progressive knowledge-to-decision alignment framework for explainable hateful meme detection. ProKDA separates explanation generation and label prediction into three sequential training stages—background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment—reducing task interference. Experiments on three public benchmarks demonstrate that ProKDA achieves state‑of‑the‑art detection performance while providing accurate, evidence‑supported explanations for moderation decisions.

By Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia
arXiv AI
Aug 20

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.

By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
arXiv AI
Aug 25

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

The paper presents a scalable approach to harmful content moderation on social media by leveraging large language models (LLMs) for few-shot, in-context learning. Experiments across multiple LLMs show that this method outperforms proprietary baselines such as Perspective and OpenAI Moderation, as well as prior few-shot learning techniques, in detecting harmful content. The study also explores the addition of visual cues like video thumbnails to assess multimodal improvements, highlighting the advantages of LLM-based moderation for dynamic and large-scale content filtering.

By Akash Bonagiri, Lucen Li, Rajvardhan Oak, Zeerak Babar, Magdalena Wojcieszak, Anshuman Chhabra