arXiv AI
Sep 2

Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching

The paper describes a new approach for two‑sided service marketplaces that replaces fixed request forms with AI‑native probabilistic matching using large language models. It introduces an autoresearch loop that generates a provider‑side preference taxonomy for each occupation, iteratively refining candidate tag sets through a six‑rubric LLM judge and a seven‑critic panel. The system also maps legacy form questions back to the new taxonomy, enabling coverage assessment and human quality assurance.

By Kartik Ravisankar, Hojat Abdolanezhad, Daniel Capo, Sang Su Lee, Shishir Dash, Vijay Anand Raghavan
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv AI
Aug 18

Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

arXiv:2608. 15877v1 Announce Type: new Abstract: Search and recommendation serve a shared discovery objective but encode intent differently.

By Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Hong Yan, Shujian Bu
arXiv Computation and Language
4d ago

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker

ZooWork-ShopRanker is a family of open e‑commerce rerankers (0.6B, 4B, and 8B) that align with human shopping preferences by using large language models as preference oracles to generate training pairs. The flagship 8B model serves as a teacher for the smaller 4B and 0.6B models, which are further refined on judged pairs. A new benchmark, ShopRank‑Bench, contains ~10,000 private‑traffic preference pairs and shows that all ZooWork models outperform the strongest open reranker baseline and their own un‑aligned versions.

By Siqiao Xue, Shuxuan Liu, Ning Hu
arXiv Computation and Language
Sep 25

Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit

The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.

By Dmitrij \.Zatuchin