CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation
arXiv:2608. 09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products.
arXiv:2606. 14715v1 Announce Type: cross Abstract: LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors.
arXiv:2608. 09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products.
arXiv:2606. 06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts.
arXiv:2606.06443v3 Announce Type: replace Abstract: Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it...
arXiv:2607. 05999v1 Announce Type: new Abstract: LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are often difficult to compare with explicit social dynamics.
The paper introduces a digital‑twin framework that simulates opinion dynamics in real Twitter networks by assigning agents attributes such as persona, emotions, centrality, stubbornness, and influence, and using Mistral‑7B to update opinions based on memory and social exposure. Validation on COVID‑19 and U.S. election 2020 datasets shows the framework reproduces opinion trajectories, reducing prediction error by over 50% compared to classical baselines, and improves structural alignment and polarization dynamics. Ablation studies reveal that agent attributes, memory, and social exposure all contribute to predictive fidelity, with agent attributes being the most critical.
The paper reports the first large‑scale empirical comparison of AI‑agent and human online communities, analyzing 73,899 Moltbook and 189,838 Reddit posts across five matched communities. It finds that Moltbook shows extreme participation inequality (Gini = 0.84 vs. 0.47) and high cross‑community author overlap (33.8% vs. 0.5%). Linguistically, AI‑generated content is emotionally flattened, more assertive than exploratory, and socially detached, leading to community‑level homogenization that is largely a structural artifact of shared authorship. At the individual level, AI agents are more identifiable than human users due to outlier stylistic profiles amplified by their extreme posting volume.
arXiv:2507.19364v3 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) into social simulation has generated considerable enthusiasm, but also raises substantial methodolo...
The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.
The study examines how four large language models (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) exhibit repetition effects in a social media simulation. Using a two‑phase within‑context design, researchers collected 336,000 ratings on truth, importance, sentiment, and interest for 100 statements across 10 feed variants and 3 replications. Results show distinct patterns: Gemma-3 displays a genuine Illusory Truth Effect, Qwen2.5 shows a mere exposure effect, GPT-5-nano shows no truth boost and mild skepticism, and Llama-3.1 shows a small truth boost with reduced evaluative dimensions.
The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
arXiv:2512. 07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies.
arXiv:2509. 00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited information.