arXiv AI

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

The study examines how users of a major dating platform respond to autonomous LLM agents that converse on their behalf. Using two large surveys, researchers built a latent-variable model showing that willingness to send and receive agent-mediated messages are highly correlated yet distinct. The findings reveal a delegation asymmetry: users are more willing to deploy their own agent than to engage with others’ agents, leading to low overall reciprocity and gender‑directional imbalances in agent interactions.

Hugging Face Trending Papers
6d ago

From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling

The paper introduces a dynamic bipartite matching framework that uses large language model (LLM) agents and contextual bandits to model decentralized, asynchronous matching processes without requiring full preference rankings. In a simulated Chinese marriage market, LLM agents evaluate local candidates while Logistic-UCB models learn reciprocal acceptance, leading to higher mutual welfare and fewer blocking pairs compared to classical Gale–Shapley. The study validates LLM-generated preferences against empirical data and demonstrates gender-differentiated acceptance patterns, supporting the use of decentralized LLM-based matching for economic simulation and computational social science.

arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv AI
Aug 17

MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

arXiv:2608. 14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims.

By Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi
arXiv AI
Aug 13

Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

arXiv:2608. 11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective.

By Alexander Liss, Nicholas Desmond, Santiago Gil Gallego
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv AI
Aug 11

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

arXiv:2608. 07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing.

By Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang