arXiv AI

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

arXiv:2607. 13998v1 Announce Type: cross Abstract: The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms.

arXiv AI
Jun 2

Ev-Trust: An Evolutionarily Stable Trust Mechanism for Decentralized LLM-Based Multi-Agent Service Economies

arXiv:2512. 16167v3 Announce Type: replace-cross Abstract: Decentralized LLM-based multi-agent service economies face three vulnerabilities that undermine traditional trust mechanisms: reduced cost of fraud, difficulty in evaluating service quality, and instability of service content.

By Jiye Wang, Shiduo Yang, Ting Qiao, Jiayu Qin, Jianbin Li, Yu Wang, Yuanhe Zhao
arXiv AI
Sep 24

Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer

The study investigates how Large Language Models (LLMs) acting as surrogate consumers are influenced by marketing pricing cues such as just‑below pricing and promotional framing. Using a tool called "Tool‑Lab" to trace information acquisition, the researchers found that when no cost is imposed, pricing cues rarely mislead LLMs, but when acquisition costs are introduced under a vague goal prompt, LLMs tend to omit important diagnostic attributes and make suboptimal choices similar to human heuristics. The findings suggest that marketing heuristics in AI‑driven shopping are shaped more by storefront information architecture than by inherent LLM limitations.

By Davood Wadi, Yu Ma
Hugging Face Trending Papers
Jul 5

Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems

The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward autonomous and interactive systems that can reason, plan, and act. This survey provides a comprehensive overview of this emerging landscape by introducing a unified taxonomy grounded in the level of autonomy and three core paradigms of agentic recommender systems: agent-assisted recommendation, agent-as-recommender, and agent-as-user-simulator.

arXiv Machine Learning
Jun 30

Persona-Trained Monte Carlo: Estimating Market-Outcome Distributions via Swarms of Persona-Conditioned Neural Policy Bots in a Limit Order Book

arXiv:2606. 29556v1 Announce Type: new Abstract: We propose Persona-Trained Monte Carlo (PTMC), a method for estimating distributions of market-outcome statistics by repeatedly simulating limit-order-book interaction among swarms of persona-conditioned neural-policy trading bots.

By Salavat Ishbulatov
Hugging Face Trending Papers
Jun 28

Persona-Trained Monte Carlo: Estimating Market-Outcome Distributions via Swarms of Persona-Conditioned Neural Policy Bots in a Limit Order Book

We propose Persona-Trained Monte Carlo (PTMC), a method for estimating distributions of market-outcome statistics by repeatedly simulating limit-order-book interaction among swarms of persona-conditioned neural-policy trading bots. Each run instantiates many bots sharing one trained policy network but conditioned on heterogeneous, individually sampled persona parameters drawn from a learned trader-heterogeneity distribution; the bots interact in a continuous double auction, and the resulting price path is one Monte Carlo sample.

arXiv AI
Sep 24

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.

By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei