arXiv AI

Personality engineering with AI agents: A new methodology for negotiation research

The article introduces personality engineering, a method that uses AI agents to precisely model negotiator personas based on established personality frameworks. It argues that AI agents, free from human limitations, can rigorously test canonical negotiation theory, which posits that success depends on balancing empathy and assertiveness. The authors propose using the interpersonal circumplex—specifically its warmth and dominance dimensions—as a foundational coordinate system for both theory testing and AI agent design.

arXiv AI
Sep 25

How Do Users Negotiate Harmful Value Conflicts with AI Companions? A Study with Minion, a Technology Probe for In-Situ Human-AI Conflict Response

The paper examines how users handle harmful value conflicts with AI companions. By analyzing 146 posts and conducting a week-long study with 22 participants using the Minion technology probe, the authors find that users blend softer and harder strategies, especially when conflicts involve Universalism and Tradition values. The study highlights that such conflicts create asymmetric responsibility, with users shouldering unilateral repair work that AI companions cannot reciprocate, suggesting a need for platform-level safeguards.

By Qing Xiao, Xianzhe Fan, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen
arXiv AI
Sep 21

Do Personality-Tuned LLMs Make Better Social Agents?

The paper examines whether fine‑tuning large language models (LLMs) with personality‑labelled data improves their ability to act as socially interactive agents. Two small open‑weight LLMs were fine‑tuned on a corpus of personality‑labelled social media posts and dialogues, and the resulting models were evaluated in various social interaction scenarios by independent LLM judges. The findings show that the fine‑tuned models do not outperform their baseline counterparts in role‑playing personalities, though they offer comparable text quality and increased linguistic diversity for the Qwen models; low inter‑rater agreement limits confidence in the results, suggesting future work should focus on training data quality and domain alignment.

By Tim Krabbe, Xiaodan Shi
arXiv AI
2d ago

Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models

The study investigates how reputation, strategy, and emotional signals influence cooperation in generative AI models using the iterated prisoner's dilemma. Non‑reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT‑4o) showed cooperation shaped by all three factors, while reasoning models (Claude 4.6, Gemini 3, GPT‑5.2) relied more on strategy and reputation, displayed reduced emotional influence, and exhibited varied end‑game behaviors. These results highlight the growing sophistication and heterogeneity of AI social behavior, suggesting the need for standardized cooperation benchmarks.

By Celso de Melo, Zishan Feng, James Hale, Kazunori Terada, Giorgio Coricelli, Jonathan Gratch
arXiv AI
Aug 26

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

AgentWorld is a simulation framework that evaluates agentic information retrieval by incorporating diverse user personalities based on the Big Five (OCEAN) traits, stateful tool-use environments, and a pass$^k$ consistency metric with structured fault classification and partial-credit scoring. It includes a risk analyzer that uses Monte‑Carlo rollouts and advanced scoring methods to quantify trajectory brittleness and attack attribution. Experiments with conversational analytics, customer‑support agents, and adversarial stress‑testing demonstrate that personality variation reveals failure modes hidden by uniform testing, such as cross‑domain leakage, contextual drift, and significant quality gaps across personas.

By Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran