arXiv AI

ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts

ArGuard is a shared task that evaluates harmful content detection in Arabic memes and LLM prompts, featuring two tracks: Track A for multimodal hate detection in memes and Track B for harmful prompt detection in Arabic LLM safety evaluation. Fifty‑eight teams registered, 35 reached the final evaluation, and 27 submitted system‑description papers, with participants experimenting with models such as AraBERT, Jais, and Qwen3‑VL. The top systems achieved macro‑F1 scores of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, with fine‑grained meme classification in A2 proving the most challenging due to sparse labels and distribution shifts.

Hugging Face Trending Papers
Jul 29

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels.

arXiv AI
Sep 18

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

The paper investigates whether large language models (LLMs) can internally detect harmful content, bypassing external guardrails that add latency and computational cost. By extracting activations from LLaMA‑3.1‑8B and training lightweight MLP probes, the authors achieve high F1 scores (99%, 83%, and 84%) on WildJailbreak, Beavertails, and AEGIS 2.0 benchmarks, rivaling much larger guard models while reducing overhead. This suggests that internal state monitoring can provide efficient safety checks for resource‑constrained, time‑critical deployments.

By Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi
arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
arXiv AI
Aug 20

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

The paper evaluates large language models for hate‑speech detection in Roman Urdu, a low‑resource language with informal spelling variations. Using the Parameter‑Efficient Fine‑Tuning technique Low‑Rank Adaptation (LoRA), the authors fine‑tune models such as Mistral, LLaMA, Falcon, and multilingual BERT on the 72,000‑comment PURUTT dataset. While zero‑shot performance yields an F1 of 0.56, fine‑tuning a small fraction of parameters boosts F1 scores above 0.93, demonstrating that PEFT offers both high accuracy and computational efficiency for low‑resource language tasks.

By Toneema Zubair, Muhammad Junaid Asif, Faisal Kamiran, Hafiz Hassan Saeed, Rana Fayyaz Ahmad
arXiv AI
Sep 25

TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification

TTLab submitted a system for the AlexandriaX-2026 Subtask 3 on Arabic machine‑translation error‑span detection and classification. The approach treats the task as token‑level classification over surface forms, using focal loss with class weighting and dialect‑specific decoding thresholds to address label imbalance. MARBERTv2, among six Arabic pre‑trained encoders, achieved the best performance, ranking third overall, though classification of rare error types remains difficult, indicating a need for data augmentation.

By Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler