From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
arXiv:2606. 03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge.
arXiv:2603. 04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- selecting the right model for each query at inference time -- has become a critical systems challenge.
arXiv:2606. 03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge.
arXiv:2608.28726v1 Announce Type: new Abstract: The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant chal...
arXiv:2608. 10714v1 Announce Type: cross Abstract: The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service provisioning that is simple to operate, scalable across independently administered domains, and agile under domain churn (i.
STR-Agent is an LLM-driven framework designed for QoS-aware routing in Low Earth Orbit satellite networks. It integrates intent perception, tool-based execution, experience accumulation, and reflection-based policy adaptation to translate natural-language service requests into adaptive routing decisions. In simulations on a Walker-Delta constellation, STR-Agent reduces end-to-end delay by up to 60% compared with DQ-Dijkstra and improves intent-understanding accuracy from 45.4% to 92.45% after fine-tuning, with the Reflection Module providing additional delay reductions.
arXiv:2510. 19366v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs).
arXiv:2603.04445v3 Announce Type: replace-cross Abstract: The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent mo...
arXiv:2605. 17106v2 Announce Type: replace-cross Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences.
The paper introduces a SAREF-compliant ontology designed to represent distributed AI workflows across edge, fog, and cloud environments. It extends the SAREF4SYST ontology with concepts for AI pipelines, executable jobs, resources, deployment constraints, and communication links, creating a unified semantic model for both AI workflows and heterogeneous infrastructures. Evaluation through smart‑grid energy service scenarios and competency questions demonstrates successful deployment, reasoning, and workload adaptation, achieving 90‑100% deployment success and sub‑80 ms orchestration times.
arXiv:2606. 22902v3 Announce Type: replace Abstract: Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all.
arXiv:2604. 26508v2 Announce Type: replace-cross Abstract: Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded platforms.
arXiv:2608. 14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison.
arXiv:2609.38697v1 Announce Type: cross Abstract: We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources...