arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
The paper introduces a dynamic bipartite matching framework that uses large language model (LLM) agents and contextual bandits to model decentralized, asynchronous matching processes without requiring full preference rankings. In a simulated Chinese marriage market, LLM agents evaluate local candidates while Logistic-UCB models learn reciprocal acceptance, leading to higher mutual welfare and fewer blocking pairs compared to classical Gale–Shapley. The study validates LLM-generated preferences against empirical data and demonstrates gender-differentiated acceptance patterns, supporting the use of decentralized LLM-based matching for economic simulation and computational social science.
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv:2602. 12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes.
By Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian
arXiv:2608.29464v1 Announce Type: cross
Abstract: Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness...
By Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
arXiv:2608. 14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims.
By Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi
arXiv:2608. 11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective.
By Alexander Liss, Nicholas Desmond, Santiago Gil Gallego
arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.
By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv:2609.24876v1 Announce Type: new
Abstract: In multi-agent social settings, model reliability varies across relationships. Beyond inferring what others will do, an agent must calibrate how confid...
By Harshil Shah, Andrew Pashea
arXiv:2608. 07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing.
By Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang
arXiv:2607. 25253v1 Announce Type: new Abstract: Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user.
By Deyao Hong, Kehan Zheng, Qian Li, Jun Zhang, Jie Jiang, Hongning Wang
arXiv:2606. 18258v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries.
By Sunnie S. Y. Kim, Margit Bowler, Leon A Gatys