arXiv Computation and Language

AI-inferred expressed well-being and collective-action discourse in climate-change campaigns on X

arXiv AI
Jul 21

Posts of Peril: Detecting Information About Hazards in Text

arXiv:2405. 17838v3 Announce Type: replace-cross Abstract: Socio-linguistic indicators of affectively-relevant phenomena, such as emotion or sentiment, are often extracted from text to better understand features of human-computer interactions, including on social media.

By Keith Burghardt, Daniel M. T. Fessler, Chyna Tang, Anne Pisor, Kristina Lerman
arXiv Computation and Language
Sep 1

Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative. "whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."

By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
arXiv Computation and Language
Sep 16

Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising

arXiv:2609.16997v1 Announce Type: new Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapid...

By Md. Samiul Alim, Mahir Shahriar Tamim, Tanvir Ahmed Khan, Sharjil Khan, Rafia Ferdous Duti, Shahriyar Zaman Ridoy, Mohammad Ali Moni
arXiv Computation and Language
6d ago

The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale

The Illinois Social Attitudes Aggregate Corpus (ISAAC) is an open, modular corpus comprising over 527 million English‑language Reddit posts from 2007 to 2023, curated for relevance to six social group distinctions—race, sexuality, age, ability, body weight, and skin tone. A multi‑step, human‑audited filtering pipeline keeps irrelevant content below 10% overall and per group, while each post receives algorithmic annotations of user home region and a suite of validated semantic labels such as moralization, sentiment, emotion, and linguistic generalization. ISAAC’s publicly available, reproducible pipeline enables cross‑category comparisons, high‑precision tracking of long‑term temporal shifts, and spatial mapping of public opinion and policy outcomes, and can be extended to new platforms, languages, and social categories via both point‑and‑click and programmatic interfaces.

By Babak Hemmatian, Sarah Hadjarab, Jessica Chen, Benedek Kurdi
arXiv AI
Aug 5

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.

By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
arXiv Machine Learning
Jun 25

Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate Discourse

arXiv:2601. 13317v2 Announce Type: replace-cross Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally different environments: paid advertising platforms host targeted, institutionally produced messaging, while public social media reflects largely organic, user-driven discussion.

By Samantha Sudhoff, Pranav Perumal, Zhaoqing Wu, Tunazzina Islam
arXiv Machine Learning
Sep 18

Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks

The paper introduces a digital‑twin framework that simulates opinion dynamics in real Twitter networks by assigning agents attributes such as persona, emotions, centrality, stubbornness, and influence, and using Mistral‑7B to update opinions based on memory and social exposure. Validation on COVID‑19 and U.S. election 2020 datasets shows the framework reproduces opinion trajectories, reducing prediction error by over 50% compared to classical baselines, and improves structural alignment and polarization dynamics. Ablation studies reveal that agent attributes, memory, and social exposure all contribute to predictive fidelity, with agent attributes being the most critical.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun, Sherali Zeadally
arXiv AI
Sep 2

Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models

The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.

By Tian Fang, Ga\"el Guibon, Davide Buscaldi