arXiv AI By Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

Read the original on arXiv AI →

arXiv:2605. 07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 16

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification

The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.

By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
arXiv AI
3d ago

RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce

RealWorldShop introduces a new benchmark for conversational shopping agents, featuring 3.28 million grounded products, structured shopping episodes, a profile‑grounded user simulator, and role‑play evaluation. Analysis reveals that existing systems generate locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially in ambiguous or multi‑intent scenarios. The authors propose REALSHOP_AGENT, a session‑control framework with explicit state management, shopping‑flow control, catalog‑grounded retrieval, and runtime guards, which consistently outperforms strong baselines on the benchmark.

By Xinwei Yang, Kelong Mao, Yudong Guo, Sulong Xu, Simiu Gu, Chen Huang, Wenqiang Lei