arXiv:2606. 00016v1 Announce Type: cross Abstract: Detecting AI-generated text is becoming increasingly challenging as modern language models approach human-level fluency and can evade detectors that rely on surface statistics or likelihood-based signals.
By Aria Nourbakhsh, Adelaide Danilov, Christoph Schommer, Salima Lamsiyah
arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.
By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
arXiv:2608. 11049v1 Announce Type: cross Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives.
By Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov
arXiv:2606. 10099v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated influence operations, motivating the need for robust detectors.
By Rafael Rivera Soto, Barry Chen, Nicholas Andrews
ProBel is a bilingual Arabic and English resource for propaganda detection that aligns binary labels, multi-label annotations for 23 propaganda techniques grouped into six categories, technique-labeled spans, and reference explanations for news sentences. The dataset supports matched binary, coarse-grained, multi-label, and span-level tasks in both languages, and the authors evaluate zero‑shot prompting, task‑specific fine‑tuning, and joint training. A single bilingual multi‑task model achieves the best overall performance, with cross‑task analysis revealing that joint classification preserves binary performance while span‑only training can weaken sentence‑level prediction, and that joint bilingual training yields the most stable results.
By Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori, Giovanni Da San Martino, Firoj Alam
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.
arXiv:2606. 04177v1 Announce Type: cross Abstract: Interpretable linguistic features offer a promising approach for explaining why a given text appears machine-generated, particularly for non-expert users.
By Yassir El Attar, Esra D\"onmez, Maximilian Maurer, Agnieszka Falenska
arXiv:2606. 04199v1 Announce Type: cross Abstract: The increasing use of large language models has raised concerns about the spread of AI-generated fake news, particularly under varying prompting strategies.
By Aya Vera-Jimenez, Samuel Jaeger, Calvin Ibenye, Dhrubajyoti Ghosh
Understanding moral values in social media text offers insight into moral judgement formation, and supervised NLP models trained on crowdsourced data have achieved strong classification performance. However, most approaches simplify the problem by aggregating multiple annotators' labels into a single "ground truth", overlooking the inherent subjectivity of the task.
Bias in natural language remains a persistent challenge in both human-written and AI-generated content, affecting domains such as journalism, education, and AI research. Most existing detection methods identify only the presence of bias, with limited support for granular detection, interpretable explanations, neutral rewriting, and openly available trained models.
LLMTrace is a new large‑scale bilingual (English and Russian) corpus designed to improve AI‑written text detection. It contains character‑level annotations that enable precise localization of AI‑generated segments, supporting both full‑text binary classification and interval detection tasks. The dataset is built from a diverse set of modern proprietary and open‑source LLMs to address gaps in existing resources, such as outdated models, limited language coverage, and lack of mixed human‑AI authorship data.
By Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Maksim Kuprashevich
arXiv:2608.22922v1 Announce Type: new
Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourc...
By Thisen Ekanayake, Nisansa de Silva