MatrAIx: Simulating the World with 8.3 Billion Persona Agents
arXiv:2608. 04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.
The paper introduces a simulation framework that uses large language model (LLM) agents conditioned on data‑driven personas to predict A/B test outcomes. These personas are built from anonymized user behavioral patterns, engagement signals, and inferred demographics, offering a more realistic population model than synthetic or rule‑based personas. The authors evaluate question design, persona data source, behavioral depth versus diversity, and population subsampling, achieving 0.75–0.90 directional accuracy on 40 real A/B tests.
arXiv:2608. 04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.
arXiv:2606. 18263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks.
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
arXiv:2609.00014v1 Announce Type: cross Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid,...
arXiv:2608. 19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems.
PersonaForge is a user‑simulation framework that generates realistic multi‑turn interactions between users and agentic systems, addressing the gap that most training data assumes single‑turn queries. It uses a four‑dimensional persona space, SOUL‑driven behavioral control calibrated to real‑user statistics, and Reverse Deep Construction from authentic seed queries to create a 6.3K‑record training set and a 138‑task benchmark called PersonaForge‑Bench across 20 professional domains. Experiments with Qwen3.5‑27B show that training with PersonaForge improves composite scores by 4.1%, especially in Task Completion (+6.0%) and Response Quality (+6.8%), while also reducing turns and tool calls, indicating more efficient interactions.
arXiv:2606. 17441v1 Announce Type: cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies.
arXiv:2602. 12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-specific preferences and latent constraints of individual users.
arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.
arXiv:2602.00685v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable an...
arXiv:2608.22438v1 Announce Type: new Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response...