arXiv AI By Alexander Liss, Nicholas Desmond, Santiago Gil Gallego

Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

Read the original on arXiv AI →

arXiv:2608. 11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 12

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.

arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
Hugging Face Trending Papers
Sep 17

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

The paper examines how large language model (LLM) based graphical user interface (GUI) agents respond to digital nudges. Using a randomized online shopping experiment with 3,600 agents across six frontier models, it finds that agents are vulnerable to both automatic and reflective nudges. The study shows that the agents’ reasoning configuration moderates these effects in opposite directions—reducing susceptibility to automatic nudges while increasing it to reflective social influence nudges—and that this redirection is systematically linked to model scale.

arXiv AI
Sep 18

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

The study examines how large language model (LLM)–based graphical user interface (GUI) agents respond to digital nudges. Using Dual‑Process Theory, researchers tested 3,600 agents across six frontier models in an online shopping experiment and found that the agents were susceptible to both automatic (Type 1) and reflective (Type 2) nudges. The agents’ reasoning configuration moderated these effects in opposite directions: extensive reasoning reduced susceptibility to automatic default nudges but increased susceptibility to reflective social‑influence nudges, with the effect systematically varying by model scale.

By Haya Halimeh, Sascha Kaltenpoth, Kevin B\"osch, Oliver M\"uller
arXiv AI
Aug 26

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent–Feedback Evolution) is a framework that lets a shared‑parameter model alternate between acting as a search agent and as a critic that provides corrective feedback. By learning when to request feedback and how to use it, CAFE trains the agent to recover from its own failures and shapes rewards both online and offline. Experiments on seven search benchmarks show that CAFE outperforms other RL‑based agents, maintains gains on out‑of‑domain tests, and reduces hallucinations, indicating that co‑evolving feedback is essential for self‑improving search agents.

By Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang