arXiv AI By Alankar Atreya, Stefan Sylvius Wagner, Devesh Batra, Robert Hankache, Cristovao Iglesias Jr, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi

Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes

Read the original on arXiv AI →

The paper presents an AI‑powered triaging agent for banks that uses large language models to conduct multi‑turn conversations, ask relevant questions, and classify customer cases for accurate routing to specialist teams. The system is integrated with policy, safety guardrails, and reasoning frameworks, and its performance is evaluated using synthetic digital twins that simulate realistic, labeled dialogues based on historical data. Results show a 30.6% increase in classification accuracy and high satisfaction from subject‑matter experts, demonstrating the effectiveness of targeted probing for scalable banking operations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 31

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

arXiv:2607. 26060v1 Announce Type: cross Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment.

By Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
arXiv AI
2d ago

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.

By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv AI
Aug 20

A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation

The paper introduces a multi‑agent platform built on CrewAI for conversational business intelligence. Five specialized agents process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol, and deliver actionable insights. The system includes a defense‑in‑depth security architecture, a query parameterization mechanism, and achieves 95.3% functional accuracy with a 24‑second mean latency, outperforming a single‑agent baseline by 22.6 percentage points in accuracy and 20.2% in quality.

By Manoj N M, Vijayakrishna S, Manjunath Srinivas, Rohit Pahan