arXiv Machine Learning

DBpedia-Enriched Company Representation for B2B Lead Recommendation

arXiv:2606. 28355v1 Announce Type: cross Abstract: Selecting which companies to approach is a central challenge in business-to-business (B2B) sales, where decisions are often based on manual research and fragmented information sources.

arXiv AI
Aug 28

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

CorporateBench (CB) is a large‑scale, human‑validated Q&A benchmark designed to evaluate large language models on enterprise‑scale document collections. It contains over 230,000 documents derived from four synthetically generated firms, each modeled with a temporally evolving knowledge base that ensures logical consistency across hundreds of thousands of documents. The benchmark tests LLMs on information extraction and knowledge‑base querying, revealing that performance degrades as input size approaches realistic corporate scales.

By Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
arXiv Machine Learning
Jun 9

Towards Personalized Bangla Book Recommendation: A Large-Scale Heterogeneous Book Graph Dataset

arXiv:2602. 12129v2 Announce Type: replace-cross Abstract: Personalized book recommendation in Bangla literature has been constrained by the lack of structured, large-scale, and publicly available datasets.

By Rahin Arefin Ahmed, Md. Anik Chowdhury, Sakil Ahmed Sheikh Reza, Devnil Bhattacharjee, Muhammad Abdullah Adnan, Julian McAuley, Nafis Sadeq
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi
arXiv AI
Aug 5

ISEE: Interactive Semantic Enrichment for Database Fields

arXiv:2608. 02604v1 Announce Type: new Abstract: LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval.

By Yuan Tian, Yiru Chen, Rakesh R. Menon, Zifan Liu, Ting Cai, Fei Wu, Anudeep Chimakurthi, Prashanthi Ramamurthy, Sridevi Aishwariya Ganesan, Kun Qian, Yunyao Li
arXiv Computation and Language
Sep 11

LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains

LLMAR is a tuning‑free recommendation framework designed for sparse, text‑rich industrial B2B domains. It transforms user behavioral history into structured semantic motives using LLM inference, employs a reflection loop to self‑correct hallucinations, and operates cost‑effectively with asynchronous batch processing. Experiments on MovieLens‑1M, Amazon Prime Pantry, and a construction risk dataset show LLMAR surpasses state‑of‑the‑art learning models, achieving up to a 54.6% nDCG@10 improvement while keeping inference costs around $1 per 1,000 users.

By Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda