arXiv:2605. 27441v2 Announce Type: replace-cross Abstract: Query understanding in large-scale industrial search systems is typically implemented as a cascade of disparate, task-specific components.
By Ping Liu, Qianqi Shen, Jianqiang Shen, Chunnan Yao, Kevin Kao, Rajat Arora, Dan Xu, Baofen Zheng, Yunxiang Ren, Benjamin Le, Ali Hooshmand, Igor Lapchuk, Juan Bottaro, Raghavan Muthuregunathan, Caleb Johnson, Liangjie Hong, Jingwei Wu, Wenjing Zhang
arXiv:2604. 26197v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents are increasingly used in real-world products, where personalized and context-aware user interactions are essential.
By Zhentao Xu, Shangjin Zhang, Emir Poyraz, Yvonne Li, Ye Jin, Xie Lu, Xiaoyang Gu, Karthik Ramgopal, Praveen Kumar Bodigutla, Xiaofeng Wang
arXiv:2608. 02604v1 Announce Type: new Abstract: LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval.
By Yuan Tian, Yiru Chen, Rakesh R. Menon, Zifan Liu, Ting Cai, Fei Wu, Anudeep Chimakurthi, Prashanthi Ramamurthy, Sridevi Aishwariya Ganesan, Kun Qian, Yunyao Li
arXiv:2609.15205v1 Announce Type: cross
Abstract: Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instruc...
By Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, Bo Li
arXiv:2608. 07023v1 Announce Type: cross Abstract: Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching.
By Emma Jouffroy, Warren Jouanneau, Marc Palyart
LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.
By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal