arXiv:2606.18389v2 Announce Type: replace
Abstract: Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated dat...
By Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin, Simon Ostermann
The paper introduces Mawqif-XT, a new Arabic benchmark dataset comprising 996 manually annotated tweets from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is labeled for stance, sentiment, and sarcasm following the Mawqif annotation scheme, and the dataset is intended as a held‑out evaluation set to test cross‑target generalization. Baseline results are provided using Arabic and multilingual transformer models as well as zero‑shot large language models, enabling reproducible evaluation alongside the original Mawqif dataset.
By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Maram Kurdi, Saad Ezzini, Asma Yamani, Ahmed Ashraf
The paper presents an agentic framework for detecting conspiratorial content in social media by inferring the speaker’s intent rather than merely identifying explicit claims. It leverages social context and adaptive tool use, demonstrating superior performance over text-only and non-agentic models on a large Hebrew tweet dataset spanning election cycles and the COVID pandemic. The study highlights the importance of context-aware, reasoning-driven approaches for accurate conspiracy detection.
By Lior Biton, Oren Tsur
arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.
By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv:2512. 14332v2 Announce Type: replace-cross Abstract: The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately.
By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
ViTOED is a new dataset for target‑oriented emotion detection in Vietnamese social media, containing 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity). The dataset uncovers Vietnamese‑specific linguistic phenomena such as implicit sources and targets and vocabulary ambiguities, and it serves as a benchmark for evaluating Vietnamese pre‑trained language models. A baseline using structured sentiment graphs shows that span detection and relation extraction remain challenging, indicating significant room for improvement in Vietnamese target‑oriented emotion detection tasks.
By Chanh Vo, Son T. Luu, Ngan Luu-Thuy Nguyen
The paper reports a comprehensive study of transformer-based NLP models for detecting check-worthy social media posts, covering data collection, preprocessing, architecture selection, fine‑tuning, testing, and implementation. It focuses on multilingual models that can process English and low‑resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, and Slovak, and compares their performance to state‑of‑the‑art baselines. The work introduces multi‑label multilingual classifiers that simultaneously identify harmful content and posts containing verifiable factual claims efficiently.
By Sebastian Kula
arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.
By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
IndicTriMix presents a new benchmark and models for token‑level language identification in tri‑language code‑mixed text involving Hindi, Gujarati, and Bengali. The authors reformulate the task as sequence labeling and fine‑tune transformer models MuRIL and XLM‑RoBERTa, evaluating them on manually annotated test sets. They also introduce two code‑mixing generation methods using parallel sentences and release the datasets and fine‑tuned models for public use.
By Pruthwik Mishra, Rudra Trivedi, Avi Patel, Ashok Urlana, Shrikant Malviya
This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity) that follow strict guidelines.
The paper presents an LLM-based framework for automatically classifying crisis levels in psychological support hotlines, addressing variability in human judgments and staffing constraints. It introduces a paralinguistic injection method that embeds non‑verbal emotional cues into transcripts, allowing the model to consider acoustic nuances. A reasoning‑enhanced training strategy encourages the model to produce diagnostic reasoning chains, which regularizes and improves classification, achieving a macro F1‑score of 0.802 and accuracy of 0.805 in 5‑fold cross‑validation.
By Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang
arXiv:2609.37891v1 Announce Type: cross
Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko