The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.
By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
arXiv:2508. 13187v4 Announce Type: replace-cross Abstract: Homelessness is a persistent social challenge, impacting millions worldwide.
By Jonathan A. Karr Jr., Benjamin F. Herbst, Matthew L. Sisk, Xueyun Li, Ting Hua, Matthew Hauenstein, Georgina Curto, Nitesh V. Chawla
The paper maps the nascent social‑science literature on large language models (LLMs) by analysing 198 curated papers and 47,719 field‑scale papers. It identifies three main domains—LLM as Social Minds, LLM Societies, and LLM‑Human Interactions—each containing 13 subcategories such as reasoning, bias, collective intelligence, and trust. The taxonomy is validated through clustering stability, author classification agreement, and topic mapping, revealing differing prominence across conference and journal venues.
By Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao
The paper introduces the Persuasion Index (PI), a taxonomy of 15 persuasion dimensions grounded in psychological and communication theories, implemented with 55 lexicon- and rule-based sub-features. PI is modular, allowing individual detectors to be swapped while preserving its theoretical framework. Evaluations on four English argumentative datasets show that PI provides a shared, lightweight feature space that captures meaningful predictive signals and reveals consistent dimension-level associations with persuasion outcomes, with variations across topics and stances.
By Liancheng Gong, Zhiyang Wang, Yiwei Xu, Julia Mendelsohn
The paper presents a knowledge graph built from Harvard Dataverse’s public data, linking 102,650 datasets to 215,985 nodes and 528,003 edges that include keywords, publications, subjects, journals, and locations. About 43,991 datasets contain geospatial metadata, and 7,654 are identified as policy‑relevant, with elections and legislatures forming the largest cluster. The authors highlight the challenge of place resolution—disconnected nodes representing the same location—and propose the graph as a testbed for AI‑driven metadata enrichment and entity resolution, noting a bias toward American city‑level data.
By Danny EBanks, Devika Jain
The paper introduces TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It combines a PIDN module that uses large language models, style transfer, and unsupervised domain adaptation to detect ideologies and filter noise, with a PIPN module that employs temporal graph neural networks to predict future ideological shifts. The authors release two large-scale datasets and validate the approach on platforms such as X and Truth Social, offering empirical insights into political polarization and ideology evolution.
The paper introduces TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It combines a PIDN module that uses large language models, style transfer, and unsupervised domain adaptation to detect ideologies and filter noise, with a PIPN module that employs temporal graph neural networks to predict future ideological shifts. The authors release two large-scale datasets and validate the approach on platforms such as X and Truth Social, offering empirical insights into political polarization and online ideology evolution.
By Yijie Xu, Chao Wang, Hui Xiong
CHRONOBERG is a temporally structured corpus of English book texts covering 250 years, curated from Project Gutenberg and enriched with temporal annotations. It enables quantification of lexical semantic change via time‑sensitive Valence‑Arousal‑Dominance analysis and the creation of historically calibrated affective lexicons. Experiments show that language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts, highlighting the need for temporally aware training and evaluation pipelines.
By Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski
The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative.
"whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."
By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
The paper introduces an open, modular AI framework that automatically detects and structures evidence of social tipping points in climate literature at the passage level. It integrates a DistilBERT boundary splitter, an iteratively augmented RoBERTa classifier, a Mistral 7B rewrite model, a LLaMA 3.2 3B rating model, and a Milvus vector store, all accessible via a Streamlit interface. Evaluation on a GPT‑4.1‑labelled benchmark and expert‑reviewed set shows the splitter outperforms competitors and the RoBERTa detector achieves high accuracy and agreement.
By Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicol\`o Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian
The paper investigates which demographic attributes large language models (LLMs) default to when annotating text without explicit demographic cues. By comparing non‑demographic, placebo‑conditioned, and demographic‑conditioned prompts on politeness and offensiveness tasks in the POPQUORN dataset, the authors find that LLMs exhibit notable gender, race, and age influences in their annotations. This contrasts with earlier studies that reported no such effects, highlighting the importance of considering demographic bias in LLM‑based annotation workflows.
By Johannes Sch\"afer, Aidan Combs, Christopher Bagdon, Jiahui Li, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie W\"uhrl, Sabine Weber, Roman Klinger
arXiv:2607. 21842v1 Announce Type: cross Abstract: Research on political polarization on social media depends on the ability to reliably measure partisanship in user-generated content.
By Fathima Ameen, Christopher G. Healey