arXiv AI By Rafa{\l} {\L}ab\k{e}dzki, Patryk Miziu{\l}a, Hubert Rutkowski, Szymon Betlewski, Cezary Depta, Szymon Janowski, Jaros{\l}aw Kochanowicz, Jan Kanty Milczek

Business Utility of Large Language Models as Exploratory Data Analysis Agents

Read the original on arXiv AI →

arXiv:2606. 00051v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research

The paper introduces DSA, an evidence‑aware orchestration framework that uses large language model agents to conduct multi‑market stock research. DSA structures the workflow into stages of evidence acquisition, context construction, model‑routed analysis, optional role and Strategy Skill reasoning, and report generation, offering both a default and an agentic profile with distinct output validation and risk safeguards. The reference implementation supports six regional markets, fifteen Strategy Skills, and multiple execution surfaces, and has passed 1,457 portable offline backend contract tests, confirming implementation conformance.

By Linsen Zhu, Yi Shi
arXiv AI
Sep 25

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.

By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv Machine Learning
Sep 16

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification

The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.

By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde