arXiv AI

CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models

arXiv:2607. 14396v1 Announce Type: new Abstract: Product catalogs are the backbone of e-commerce sites, yet a large number of structured attributes (SAs) -- such as material, color, and shape -- often have missing values.

arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
Hugging Face Trending Papers
Aug 13

Latent On-Policy Self-Distillation

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.

arXiv AI
Jun 9

Exploring Autonomous Agentic Data Engineering for Model Specialization

arXiv:2605. 30407v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.

By Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng
arXiv AI
Aug 28

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

PACEShop introduces a new evaluation framework, PACE, for shopping assistants that emphasizes personalized, actionable, compositional, and evidence‑grounded responses. The benchmark dataset contains 22,625 records with structured personas, auditable evidence pools, and detailed defect annotations, while PACEJudge offers a training‑free protocol for assessing these dimensions. Experiments demonstrate that generic judges miss key diagnostic fields, whereas PACEJudge improves evaluation across persona alignment, cross‑component consistency, grounding, and defect localization without retraining.

By Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu
arXiv Computation and Language
Sep 1

LLP: LLM-Based Product Pricing in E-commerce

The paper introduces LLP, a Large Language Model–based generative framework for pricing second‑hand products on consumer‑to‑consumer platforms. LLP retrieves similar items to capture market dynamics, then uses LLMs to generate price suggestions, refined through supervised fine‑tuning and group relative policy optimization. A confidence‑based filter rejects unreliable predictions, and experiments show LLP outperforms prior methods, achieving higher static adoption rates when deployed on Xianyu.

By Hairu Wang, Sheng You, Qiheng Zhang, Xike Xie, Shuguang Han, Yuchen Wu, Fei Huang, Jufeng Chen