arXiv Machine Learning

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

arXiv:2606. 06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts.

arXiv Computation and Language
Sep 18

Finding Common Ground: Graded Communal Knowledge in Bluesky Starter Packs

The paper investigates how shared community affiliations, measured via Bluesky starter packs, correlate with common ground between users. By analyzing 191,648 user pairs, it finds that lexical similarity—used as a proxy for common ground—increases monotonically with the number of shared starter packs, especially when those packs represent distinct topical communities. The study also shows that this effect is independent of network proximity, indicating that community membership is a distinct, measurable carrier of common ground.

By Sagar Kumar, Lawrence Swaminathan Xavier Prince, Julia Mendelsohn, Brooke Foucault Welles, Nicholas W. Landry
Hugging Face Trending Papers
Jul 29

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. We propose IRIS, a framework that learns dynamic user personas directly from implicit interaction streams by extracting behavioral signals from everyday conversations and iteratively refining persona representations through a prediction-driven closed loop without requiring explicit feedback.

arXiv AI
Sep 21

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

The paper introduces the COmmunity-centered Peer Engaged Support (COPES) dataset and a three‑axis evaluation framework to gauge how well Large Language Models (LLMs) align with community perspectives on mental‑health support queries. Experiments show that fine‑tuning LLMs on COPES improves strategy alignment and emotion‑tone alignment by over 50% for general‑purpose models, yet these gains are uneven across subreddits and coping strategies. The study also finds that post‑training shifts the model’s recommendations toward problem‑focused advice while reducing emotion‑focused responses, indicating persistent disparities in performance across different communities and needs.

By Mohit Chandra, Nabin Kim, Eli Min, Aamogh Sawant, Tanmay Sutar, Munmun De Choudhury
arXiv Machine Learning
Jul 30

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

arXiv:2607. 26473v1 Announce Type: new Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings.

By Haifeng Wu
arXiv Computation and Language
Sep 18

Social Simulacra in the Wild: AI Agent Communities on Moltbook

The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.

By Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, Koustuv Saha
arXiv Computation and Language
Sep 14

PACIFIC: Can LLMs Discern the Psychometric Traits Influencing Your Preferences? Personality-Driven Preference Alignment in LLMs

PACIFIC is a framework that aligns large language model responses with user preferences by leveraging stable Big‑Five personality traits as a latent signal. The authors built a 1,200‑pair dataset covering diverse domains and trait directions, and found that trait‑aligned contexts enable LLMs to achieve near‑ceiling accuracy (up to 99%) in personalized QA. They also introduced a persona‑aware contrastive retriever (PiRAG) that improves label‑free accuracy from 30% to 43% over standard semantic retrieval, highlighting retrieval as the main bottleneck.

By Tianyu Zhao, Siqi Li, Yasser Shoukry, Salma Elmalaki