Strategic Buying Agents
arXiv:2607. 04708v1 Announce Type: cross Abstract: Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf.
Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf. We study the design of such strategic buying agents, which must decide when to purchase within a finite shopping window, translating price observations, the remaining time horizon, and beliefs about future price changes into a purchase policy.
arXiv:2607. 04708v1 Announce Type: cross Abstract: Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf.
arXiv:2601. 01279v3 Announce Type: replace-cross Abstract: When competing sellers delegate pricing to a shared AI model, such as a large language model, correlated recommendations combined with performance-driven updates aggregating seller feedback raise a key question: can standard AI deployment practices inadvertently produce supracompetitive pricing?
The paper proposes an adversarial reinforcement‑learning framework for market making that incorporates Hawkes‑process driven order arrivals and trade‑induced price impact, addressing limitations of prior Poisson‑based models. An LSTM module captures temporal dependencies in recent observations to handle increased non‑stationarity, and the authors analyze equilibrium properties and introduce a robustness evaluation protocol focused on the left tail of returns. Experiments across diverse market regimes demonstrate that the method improves left‑tail performance, especially under strong Hawkes excitation and moderate price impact, without relying on a terminal inventory bias.
arXiv:2512. 09850v2 Announce Type: replace Abstract: We introduce Conformal Bandits, a novel framework integrating Conformal Prediction (CP) into bandit problems, a classic paradigm for sequential decision-making under uncertainty.
arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).
arXiv:2602. 17086v2 Announce Type: replace-cross Abstract: Dynamic decision-making under model uncertainty is central to many economic environments, yet existing bandit and reinforcement learning algorithms rely on the assumption of correct model specification.
arXiv:2605. 00369v4 Announce Type: replace-cross Abstract: We study how large language models can be used to generate inventory policies in online settings with non-stationary demand.
arXiv:2603. 16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
The paper introduces a new approach to safety in contextual bandits with continuous actions by enforcing high‑probability constraints on the realized cost rather than on its expectation. It proposes the High‑Probability Constrained UCB algorithm, which balances optimistic reward exploration with pessimistic safety estimation. The authors provide theoretical regret guarantees for linear models and extend the analysis to general function classes, demonstrating experimentally that realized‑cost constraints significantly reduce safety violations compared to expected‑cost baselines.
arXiv:2606. 15862v1 Announce Type: new Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
The study investigates how Large Language Models (LLMs) acting as surrogate consumers are influenced by marketing pricing cues such as just‑below pricing and promotional framing. Using a tool called "Tool‑Lab" to trace information acquisition, the researchers found that when no cost is imposed, pricing cues rarely mislead LLMs, but when acquisition costs are introduced under a vague goal prompt, LLMs tend to omit important diagnostic attributes and make suboptimal choices similar to human heuristics. The findings suggest that marketing heuristics in AI‑driven shopping are shaped more by storefront information architecture than by inherent LLM limitations.
arXiv:2607. 24115v1 Announce Type: cross Abstract: We study the contextual dynamic pricing problem under non-stationarity, where a firm sells products to $T$ sequentially arriving consumers that behave according to an unknown demand model that can change over time.