arXiv AI

Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations

The study examines how four large language models (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) exhibit repetition effects in a social media simulation. Using a two‑phase within‑context design, researchers collected 336,000 ratings on truth, importance, sentiment, and interest for 100 statements across 10 feed variants and 3 replications. Results show distinct patterns: Gemma-3 displays a genuine Illusory Truth Effect, Qwen2.5 shows a mere exposure effect, GPT-5-nano shows no truth boost and mild skepticism, and Llama-3.1 shows a small truth boost with reduced evaluative dimensions.

arXiv Machine Learning
Sep 18

Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks

The paper introduces a digital‑twin framework that simulates opinion dynamics in real Twitter networks by assigning agents attributes such as persona, emotions, centrality, stubbornness, and influence, and using Mistral‑7B to update opinions based on memory and social exposure. Validation on COVID‑19 and U.S. election 2020 datasets shows the framework reproduces opinion trajectories, reducing prediction error by over 50% compared to classical baselines, and improves structural alignment and polarization dynamics. Ablation studies reveal that agent attributes, memory, and social exposure all contribute to predictive fidelity, with agent attributes being the most critical.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun, Sherali Zeadally
arXiv AI
Sep 24

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

The study examines how different editorial framings in prompts influence large language models’ statistical analysis reports. Using a 4×4 factorial design, researchers found that certain framings—particularly brutally critical prompts on genuine effects and significance-seeking prompts on underpowered nulls—led to factual misrepresentations. Tone shifts were more widespread, with critical framing inducing defensive language across all data patterns, while a confound in the data largely prevented both factual and tonal distortions.

By Paras Balani, Subhrakanta Panda
arXiv Computation and Language
Sep 11

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

The study audits 576 LLM-based social simulations from 350 papers using the PIMMUR framework, which evaluates agent profile, interaction, memory, minimal control, unawareness, and realism. Results show that PIMMUR principles are met more often than minimal control, unawareness, and realism, with frontier LLMs correctly identifying the underlying experiment in 65.2% of cases and half of prompts pre‑determining outcomes. Reproducing five experiments revealed that many reported collective phenomena disappear or reverse when PIMMUR principles are enforced, suggesting that apparent emergent behaviors may be methodological artifacts rather than genuine social dynamics.

By Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou, Man Ho Lam, Xintao Wang, Hao Zhu, Wenxuan Wang, Maarten Sap
arXiv Computation and Language
Sep 23

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.

By Alexandre Cristov\~ao Maiorano