arXiv:2606. 06744v1 Announce Type: new Abstract: Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation.
By Haijing Zong, Yancheng Liang, Boyang Zhou, Natasha Jaques
arXiv:2506. 03802v2 Announce Type: replace Abstract: We introduce a learning problem in a generalized two-sided matching market, where agents select actions to interact with their match.
By Andreas Athanasopoulos, Christos Dimitrakakis
arXiv:2605. 01961v2 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents.
By Maheed H. Ahmed, Mahsa Ghasemi
arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
By Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, Max Pellert
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
By Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
arXiv:2608. 12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes.
By Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2512. 04988v2 Announce Type: replace-cross Abstract: Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agentic swarms.
By Christopher Chiu, Simpson Zhang, Mihaela van der Schaar
arXiv:2606. 19883v1 Announce Type: new Abstract: We study a multi-agent multi-armed bandit problem in the competitive setup with two-sided matching markets under a human centric decision making model.
By Ananya Kunisetty, Avishek Ghosh
The study examines how users of a major dating platform respond to autonomous LLM agents that converse on their behalf. Using two large surveys, researchers built a latent-variable model showing that willingness to send and receive agent-mediated messages are highly correlated yet distinct. The findings reveal a delegation asymmetry: users are more willing to deploy their own agent than to engage with others’ agents, leading to low overall reciprocity and gender‑directional imbalances in agent interactions.
By Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
The paper introduces “SNSW-Alg”, an algorithm that finds a stable matching maximizing Nash social welfare in the stable marriage problem. It runs in ×O(n^4) time and balances equity while maintaining stability. Experiments across various preference distributions show significant fairness gains with minimal impact on regret, egalitarian criterion, and sex equality, and the resulting matchings are statistically Pareto-undominated by other fairness-based stable matchings.
By Parth Desai, Rasheed M, Ganesh Ghalme, Sujit Gujar