The Illinois Social Attitudes Aggregate Corpus (ISAAC) is an open, modular corpus comprising over 527 million English‑language Reddit posts from 2007 to 2023, curated for relevance to six social group distinctions—race, sexuality, age, ability, body weight, and skin tone. A multi‑step, human‑audited filtering pipeline keeps irrelevant content below 10% overall and per group, while each post receives algorithmic annotations of user home region and a suite of validated semantic labels such as moralization, sentiment, emotion, and linguistic generalization. ISAAC’s publicly available, reproducible pipeline enables cross‑category comparisons, high‑precision tracking of long‑term temporal shifts, and spatial mapping of public opinion and policy outcomes, and can be extended to new platforms, languages, and social categories via both point‑and‑click and programmatic interfaces.
By Babak Hemmatian, Sarah Hadjarab, Jessica Chen, Benedek Kurdi
The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.
By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
arXiv:2609.24574v1 Announce Type: new
Abstract: Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the...
By Hazem Ibrahim, Yasir Zaki
The paper introduces WSF-ARG+, a new dataset that pairs hate speech with check‑worthiness annotations, and presents an LLM‑in‑the‑loop framework to streamline the annotation process. Experiments with 12 open‑weight large language models demonstrate that the framework cuts human effort while maintaining annotation quality. The study also shows that incorporating check‑worthiness labels improves hate‑speech detection performance, boosting macro‑F1 scores for large models by up to 0.213 and averaging 0.154 across models.
By Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin
The paper introduces a diagnostic tool for distinguishing the use of misogynistic slurs from their mention in counter‑speech within code‑mixed Hinglish. It identifies evaluation artifacts in existing corpora, releases a 416‑item minimal‑pair contrast set that decorrelates slur presence and gendered register from labels, and proposes a pair‑consistency metric to assess model performance. Experiments show that even strong baselines struggle to consistently label counter‑speech pairs, while a large language model achieves perfect scores, indicating the benchmark measures genuine capability rather than exploitation of artifacts.
By Ashanvi Yadav, Shubham Bhardwaj
arXiv:2502. 10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain.
By Ezinne Nwankwo, Lauri Goldkind, Angela Zhou