Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
The paper examines how large language model (LLM) based graphical user interface (GUI) agents respond to digital nudges. Using a randomized online shopping experiment with 3,600 agents across six frontier models, it finds that agents are vulnerable to both automatic and reflective nudges. The study shows that the agents’ reasoning configuration moderates these effects in opposite directions—reducing susceptibility to automatic nudges while increasing it to reflective social influence nudges—and that this redirection is systematically linked to model scale.
The study examines how large language model (LLM)–based graphical user interface (GUI) agents respond to digital nudges. Using Dual‑Process Theory, researchers tested 3,600 agents across six frontier models in an online shopping experiment and found that the agents were susceptible to both automatic (Type 1) and reflective (Type 2) nudges. The agents’ reasoning configuration moderated these effects in opposite directions: extensive reasoning reduced susceptibility to automatic default nudges but increased susceptibility to reflective social‑influence nudges, with the effect systematically varying by model scale.
By Haya Halimeh, Sascha Kaltenpoth, Kevin B\"osch, Oliver M\"uller
CAFE (Coupled Agent–Feedback Evolution) is a framework that lets a shared‑parameter model alternate between acting as a search agent and as a critic that provides corrective feedback. By learning when to request feedback and how to use it, CAFE trains the agent to recover from its own failures and shapes rewards both online and offline. Experiments on seven search benchmarks show that CAFE outperforms other RL‑based agents, maintains gains on out‑of‑domain tests, and reduces hallucinations, indicating that co‑evolving feedback is essential for self‑improving search agents.
By Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Tre...
arXiv:2606. 29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states.
By Minghui Ma, Bin Guo, Han Wang, Mengqi Chen, Jingqi Liu, Yan Liu, Zhiwen Yu
The paper introduces a new multi‑agent micro‑benchmark called Delay‑of‑Gratification, modeled after the Stanford marshmallow experiment, to evaluate large language models (LLMs) in long‑horizon, multi‑turn interactions. In the benchmark, ReAct agents use a per‑step “raise a question” tool under various constraints—social context (broadcast vs. isolated), persona traits (age, hedonic drive), and tool‑use policy (mandatory vs. optional). Across 19,200 trajectories, the study finds that most agents exhibit an early impulse to “eat,” only 75.9% persist to the end, and factors such as isolation and hedonic drive significantly influence survival and questioning behavior, with ablations showing that removing hedonic drive and age can improve completion rates.
By Olga Manakina, Igor Bogdanov, Chung-Horng Lung
Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of conversation, a process-oriented criterion that neither surface-level text metrics (BLEU/ROUGE) nor single-score LLM judgments can capture.
arXiv:2608. 01425v1 Announce Type: cross Abstract: Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.
By Yi Mao, Andrew Perrault
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan