arXiv:2607. 08782v1 Announce Type: cross Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models.
By Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li, Fahao Chen, Haodong Wang, Jian Lin, Song Guo
arXiv:2609.37236v1 Announce Type: new
Abstract: An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested....
By Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen
The paper presents a decision‑support system that enhances retrieval‑augmented generation (RAG) for customer contact centers by first identifying customer questions in real time. If a query matches a frequently asked question (FAQ), the system retrieves the answer directly from the FAQ database; otherwise it generates an answer via RAG, delivering responses to agents within two seconds. The approach reduces manual query formulation, lowers average handling times, and cuts operational costs, and it includes an automated workflow that uses LLMs to extract FAQs from historical transcripts when none are predefined.
By Garima Agrawal, Sashank Gummuluri, Cosimo Spera
arXiv:2406. 06855v3 Announce Type: replace-cross Abstract: To leverage prediction models to make optimal scheduling decisions in service systems, we must understand how predictive errors impact congestion due to externalities on the delay of other jobs.
By Jiung Lee, Hongseok Namkoong, Yibo Zeng
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --...
arXiv:2609.37588v1 Announce Type: new
Abstract: Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the...
By T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan