Hugging Face Trending Papers

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention.

arXiv AI
Aug 25

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

AraDetox is a newly released multi-dialect Arabic detoxification dataset containing 10,500 harmful social‑media posts and 84,000 detoxified rewrites generated by GPT‑5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Human evaluation and automatic analyses confirm that the rewrites effectively remove harmful language while preserving meaning, lexical change, and dialectal style. The dataset is publicly available to support future research in Arabic detoxification, safe text generation, and multi‑dialect NLP.

By Mo El-Haj
Hugging Face Trending Papers
Sep 2

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

MultiGhostBench is a multilingual benchmark for authorship attribution of long‑form text generated by large language models. It contains 928 books produced by five recent LLMs in six languages and three scripts, each averaging about 59,000 words, and is designed to test attribution methods under domain, author, and language shifts. Experiments show that no single attribution method dominates across all settings, with performance generally dropping under distribution shifts, while transformer‑based detectors retain generator information across languages but vary in transfer effectiveness, and statistical/fingerprint detectors are more language‑dependent.

arXiv AI
Sep 3

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

MultiGhostBench is a multilingual benchmark for authorship attribution of long-form text generated by large language models. It contains 928 books produced by five recent LLMs in six languages and three scripts, each averaging about 59,000 words, and is designed to test attribution methods under domain, author, and language shifts. Experiments show that no single attribution method dominates across all settings, with performance generally dropping under distribution shifts, and that transformer-based detectors retain generator information across languages while statistical and fingerprint-based detectors are more language‑dependent.

By Matteo Greco, Anudeex Shetty, Andrea Tagarelli, Jey Han Lau
arXiv Computation and Language
Sep 7

Multilingual Models for Check-Worthy Social Media Posts Detection

The paper reports a comprehensive study of transformer-based NLP models for detecting check-worthy social media posts, covering data collection, preprocessing, architecture selection, fine‑tuning, testing, and implementation. It focuses on multilingual models that can process English and low‑resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, and Slovak, and compares their performance to state‑of‑the‑art baselines. The work introduces multi‑label multilingual classifiers that simultaneously identify harmful content and posts containing verifiable factual claims efficiently.

By Sebastian Kula
arXiv AI
Sep 10

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

The paper introduces a pipeline for creating a high‑quality European Portuguese (PT‑PT) web corpus, drawing from 411 TB of raw data from Arquivo.pt. It adds a novel post‑scraping step that removes boilerplate and duplicate lines before filtering, boosting the final document yield by 19.04%. The pipeline also incorporates language identification, weighted fuzzy deduplication, and neural quality classification to produce a clean, representative dataset suitable for large‑language‑model pre‑training.

By Gon\c{c}alo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simpl\'icio, Diogo Tavares, David Semedo, Daniel Gomes, Jo\~ao Magalh\~aes
arXiv AI
Jun 12

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508. 01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult.

By Lucio La Cava, Dominik Macko, R\'obert M\'oro, Ivan Srba, Andrea Tagarelli
arXiv AI
Aug 20

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

The paper evaluates large language models for hate‑speech detection in Roman Urdu, a low‑resource language with informal spelling variations. Using the Parameter‑Efficient Fine‑Tuning technique Low‑Rank Adaptation (LoRA), the authors fine‑tune models such as Mistral, LLaMA, Falcon, and multilingual BERT on the 72,000‑comment PURUTT dataset. While zero‑shot performance yields an F1 of 0.56, fine‑tuning a small fraction of parameters boosts F1 scores above 0.93, demonstrating that PEFT offers both high accuracy and computational efficiency for low‑resource language tasks.

By Toneema Zubair, Muhammad Junaid Asif, Faisal Kamiran, Hafiz Hassan Saeed, Rana Fayyaz Ahmad