Model in Distress: Sentiment Analysis on French Synthetic Social Media
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606.18389v2 Announce Type: replace Abstract: Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated dat...
The paper introduces Mawqif-XT, a new Arabic benchmark dataset comprising 996 manually annotated tweets from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is labeled for stance, sentiment, and sarcasm following the Mawqif annotation scheme, and the dataset is intended as a held‑out evaluation set to test cross‑target generalization. Baseline results are provided using Arabic and multilingual transformer models as well as zero‑shot large language models, enabling reproducible evaluation alongside the original Mawqif dataset.
The paper presents an agentic framework for detecting conspiratorial content in social media by inferring the speaker’s intent rather than merely identifying explicit claims. It leverages social context and adaptive tool use, demonstrating superior performance over text-only and non-agentic models on a large Hebrew tweet dataset spanning election cycles and the COVID pandemic. The study highlights the importance of context-aware, reasoning-driven approaches for accurate conspiracy detection.
arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.
arXiv:2512. 14332v2 Announce Type: replace-cross Abstract: The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately.
ViTOED is a new dataset for target‑oriented emotion detection in Vietnamese social media, containing 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity). The dataset uncovers Vietnamese‑specific linguistic phenomena such as implicit sources and targets and vocabulary ambiguities, and it serves as a benchmark for evaluating Vietnamese pre‑trained language models. A baseline using structured sentiment graphs shows that span detection and relation extraction remain challenging, indicating significant room for improvement in Vietnamese target‑oriented emotion detection tasks.