arXiv AI By Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

Read the original on arXiv AI →

The study examines how users of a major dating platform respond to autonomous LLM agents that converse on their behalf. Using two large surveys, researchers built a latent-variable model showing that willingness to send and receive agent-mediated messages are highly correlated yet distinct. The findings reveal a delegation asymmetry: users are more willing to deploy their own agent than to engage with others’ agents, leading to low overall reciprocity and gender‑directional imbalances in agent interactions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
5d ago

From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling

The paper introduces a dynamic bipartite matching framework that uses large language model (LLM) agents and contextual bandits to model decentralized, asynchronous matching processes without requiring full preference rankings. In a simulated Chinese marriage market, LLM agents evaluate local candidates while Logistic-UCB models learn reciprocal acceptance, leading to higher mutual welfare and fewer blocking pairs compared to classical Gale–Shapley. The study validates LLM-generated preferences against empirical data and demonstrates gender-differentiated acceptance patterns, supporting the use of decentralized LLM-based matching for economic simulation and computational social science.

arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv AI
Aug 17

MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

arXiv:2608. 14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims.

By Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi