arXiv Machine Learning

Positional task conditioning for scalable defect detection across product families in large product catalogs

The paper presents a method called Positional Task Conditioning (PTC) to improve defect detection in large product catalogs. By breaking detection into focused sub‑tasks and reinforcing task identity at prompt boundaries, PTC reduces context length and isolates error types, boosting F1 scores from 52% to 87%. The approach outperforms rationale‑based distillation across multiple models, achieving near‑state‑of‑the‑art performance at up to 98% lower cost and is deployed in several countries handling over 10 million product families.

arXiv Machine Learning
Aug 11

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

arXiv:2608. 07770v1 Announce Type: new Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear.

By Mike Szklarzewski, CJ George, Gavin Smithson, Christopher Stokes, Dakota Fulp, William M. Jones, Benjamin Wynn, Alexander Ur, Agit Yesiloz, Clint Kallenbach, Mark Swartz, Nathan DeBardeleben, Sharmistha Chakrabarti
arXiv AI
Aug 28

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

PACEShop introduces a new evaluation framework, PACE, for shopping assistants that emphasizes personalized, actionable, compositional, and evidence‑grounded responses. The benchmark dataset contains 22,625 records with structured personas, auditable evidence pools, and detailed defect annotations, while PACEJudge offers a training‑free protocol for assessing these dimensions. Experiments demonstrate that generic judges miss key diagnostic fields, whereas PACEJudge improves evaluation across persona alignment, cross‑component consistency, grounding, and defect localization without retraining.

By Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu
arXiv AI
Jun 9

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

arXiv:2606. 07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding.

By Zekai Zhang, Jinglin Zhang, Qinghui Chen, Gang Li, Da Chen, Shuainan Jing, He Wang, Dagang Li, Cong Liu, Cong Bai, Shengyong Chen
arXiv Machine Learning
Sep 3

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

The paper investigates how combining soft prompts via task arithmetic can reduce reliance on confounding variables in classification models. It introduces Hybrid Prompt Arithmetic (HyPA), which merges task prompts with linearized confounder prompts to counteract spurious correlations. Experiments across multiple benchmarks show that HyPA consistently improves the robustness‑performance trade‑off under distribution shift, and analysis of hidden representations suggests it mitigates confounding by diminishing the influence of confounder signals.

By Zhecheng Sheng, Yongsen Tan, Xiruo Ding, Trevor Cohen, Serguei Pakhomov
arXiv AI
Jul 17

CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models

arXiv:2607. 14396v1 Announce Type: new Abstract: Product catalogs are the backbone of e-commerce sites, yet a large number of structured attributes (SAs) -- such as material, color, and shape -- often have missing values.

By Zhu Cheng (Xuan), Zhenming Wang (Xuan), Yu (Xuan), Tang, Dan Liu, Bryan Zhang, Athanasios N. Nikolakopoulos, Pranav Souri Itabada, Jing Zhang, Chih-Chi Chou, Peng Gao, Fatemeh Mansoori, Bharat Bojja, Sarath Chander, Sameer Thombare, Umit Batur, Tarik Arici
Hugging Face Trending Papers
Jun 1

Monitoring Agentic Systems Before They're Reliable

Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape. At this maturity level, task-level error detection may be infeasible: structural failure modes mask the signal that task-level monitors are designed to detect.