EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
arXiv:2605. 12887v2 Announce Type: replace-cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes.
TRACE is a new framework that uses agentic Large Language Models to automatically enrich e-commerce product catalogs with missing or buried attributes. It employs a ScoutAgent to gather multimodal evidence from merchant catalogs, syndicated feeds, and web search, and a JudgeAgent to verify and publish the proposed attribute values. In offline evaluation, TRACE achieved 98.2% accuracy with 74.7% coverage, and in production it increased enrichment coverage by 90.4% and boosted checkout conversion by 0.48%.
arXiv:2605. 12887v2 Announce Type: replace-cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes.
arXiv:2606. 17698v1 Announce Type: new Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked.
arXiv:2609.06027v1 Announce Type: cross Abstract: Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. E...
arXiv:2605.23916v2 Announce Type: replace-cross Abstract: AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descrip...
arXiv:2605. 21347v3 Announce Type: replace Abstract: Diagnosing failures in LLM agents remains largely manual.
The paper introduces a scalable product‑linking system that uses a retrieve‑then‑match cascade. First, a lightweight text cross‑encoder auto‑resolves the majority of merchant‑catalog product pairs with high precision, while an agentic multimodal vision‑language model handles the remaining ambiguous cases by inspecting images and performing web searches. This approach balances computational cost and accuracy, improving overall link coverage from 68% to 77% without requiring fine‑tuning of the agent.
The paper introduces Agentic Share-of-Search (ASoS), a multi‑agent AI system designed to aid sellers in competitive decision‑making within large‑language‑model (LLM) mediated e‑commerce. It automates competitive visibility measurement and root‑cause diagnosis by deploying query agents on leading AI platforms and employing a ReAct‑based diagnostic agent to suggest prioritized merchandising actions. A 100‑trial ablation study demonstrates the prototype’s effectiveness, recovering the ablated signal in 39% of trials (95% CI: 30.0%‑48.8%) and 63.9% in high‑correlation cases, outperforming chance by 5.5×.
arXiv:2607. 14396v1 Announce Type: new Abstract: Product catalogs are the backbone of e-commerce sites, yet a large number of structured attributes (SAs) -- such as material, color, and shape -- often have missing values.
The paper describes a new approach for two‑sided service marketplaces that replaces fixed request forms with AI‑native probabilistic matching using large language models. It introduces an autoresearch loop that generates a provider‑side preference taxonomy for each occupation, iteratively refining candidate tag sets through a six‑rubric LLM judge and a seven‑critic panel. The system also maps legacy form questions back to the new taxonomy, enabling coverage assessment and human quality assurance.
arXiv:2608.30023v1 Announce Type: cross Abstract: Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying...
The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.
PACEShop introduces a new evaluation framework, PACE, for shopping assistants that emphasizes personalized, actionable, compositional, and evidence‑grounded responses. The benchmark dataset contains 22,625 records with structured personas, auditable evidence pools, and detailed defect annotations, while PACEJudge offers a training‑free protocol for assessing these dimensions. Experiments demonstrate that generic judges miss key diagnostic fields, whereas PACEJudge improves evaluation across persona alignment, cross‑component consistency, grounding, and defect localization without retraining.