Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
arXiv:2607. 08782v1 Announce Type: cross Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models.
arXiv:2606. 19759v1 Announce Type: new Abstract: As individuals turn to the Internet to find answers to questions they may have, several Question Answering (QA) forums have evolved, where users knowledgeable in certain topics can contribute their expertise to answering these requests for information.
arXiv:2607. 08782v1 Announce Type: cross Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models.
arXiv:2609.37236v1 Announce Type: new Abstract: An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested....
The paper presents a decision‑support system that enhances retrieval‑augmented generation (RAG) for customer contact centers by first identifying customer questions in real time. If a query matches a frequently asked question (FAQ), the system retrieves the answer directly from the FAQ database; otherwise it generates an answer via RAG, delivering responses to agents within two seconds. The approach reduces manual query formulation, lowers average handling times, and cuts operational costs, and it includes an automated workflow that uses LLMs to extract FAQs from historical transcripts when none are predefined.
arXiv:2406. 06855v3 Announce Type: replace-cross Abstract: To leverage prediction models to make optimal scheduling decisions in service systems, we must understand how predictive errors impact congestion due to externalities on the delay of other jobs.
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --...
arXiv:2609.37588v1 Announce Type: new Abstract: Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the...
arXiv:2606. 07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end.
arXiv:2606. 03014v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs.
The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.
arXiv:2608. 07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc.
arXiv:2608.22266v1 Announce Type: new Abstract: In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized assistants. A c...
arXiv:2606. 00809v1 Announce Type: new Abstract: Many real-world conversational settings for knowledge discovery, including podcasts, hiring screens, and marketplaces, require a purpose-driven understanding of a person.