Hugging Face Trending Papers

Partner-Specific Affective Precision in Social Active Inference

arXiv AI
Sep 25

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.

By Igor Bogdanov, Olga Manakina, Chung-Horng Lung
arXiv AI
Aug 19

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

The study examines how users of a major dating platform respond to autonomous LLM agents that converse on their behalf. Using two large surveys, researchers built a latent-variable model showing that willingness to send and receive agent-mediated messages are highly correlated yet distinct. The findings reveal a delegation asymmetry: users are more willing to deploy their own agent than to engage with others’ agents, leading to low overall reciprocity and gender‑directional imbalances in agent interactions.

By Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov
arXiv AI
Jul 7

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

arXiv:2604. 15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings.

By Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin
arXiv AI
2d ago

Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models

The study investigates how reputation, strategy, and emotional signals influence cooperation in generative AI models using the iterated prisoner's dilemma. Non‑reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT‑4o) showed cooperation shaped by all three factors, while reasoning models (Claude 4.6, Gemini 3, GPT‑5.2) relied more on strategy and reputation, displayed reduced emotional influence, and exhibited varied end‑game behaviors. These results highlight the growing sophistication and heterogeneity of AI social behavior, suggesting the need for standardized cooperation benchmarks.

By Celso de Melo, Zishan Feng, James Hale, Kazunori Terada, Giorgio Coricelli, Jonathan Gratch
arXiv Computation and Language
Sep 15

Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

The paper introduces the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset, comprising 2,400 prompts across five relationship themes, to study how query formulation affects sycophancy in large language models. Using the ELEPHANT framework, the authors evaluated GPT‑5 Mini and Gemini 3 Flash, finding that grammatical mood alone does not drive sycophantic behavior, whereas perspective‑driven framing does, with models increasingly accepting user premises over successive turns. Gemini 3 Flash showed smaller increases in moral sycophancy than GPT‑5 Mini, indicating greater resistance to reinforcing ethically problematic positions.

By Helena Choi, Edric Castel Hao, Karl Bautista, Francis Gabriel Magleo, Renzo Panti, Danielle Beatrice Olalia