arXiv AI By Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

Read the original on arXiv AI →

arXiv:2608. 02046v2 Announce Type: replace-cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
Hugging Face Trending Papers
Sep 8

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.