arXiv AI

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

arXiv:2603. 04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- selecting the right model for each query at inference time -- has become a critical systems challenge.

arXiv AI
Aug 12

Conversational Orchestration for Organic 6G

arXiv:2608. 10714v1 Announce Type: cross Abstract: The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service provisioning that is simple to operate, scalable across independently administered domains, and agile under domain churn (i.

By Masoud Shokrnezhad, Tarik Taleb
arXiv AI
Sep 18

STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks

STR-Agent is an LLM-driven framework designed for QoS-aware routing in Low Earth Orbit satellite networks. It integrates intent perception, tool-based execution, experience accumulation, and reflection-based policy adaptation to translate natural-language service requests into adaptive routing decisions. In simulations on a Walker-Delta constellation, STR-Agent reduces end-to-end delay by up to 60% compared with DQ-Dijkstra and improves intent-understanding accuracy from 45.4% to 92.45% after fine-tuning, with the Reflection Module providing additional delay reductions.

By Bowen Lu, Mugen Peng, Yaohua Sun, Hongyu Wang, Kerui Guo, Wenjia Xu
arXiv AI
Aug 28

SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum

The paper introduces a SAREF-compliant ontology designed to represent distributed AI workflows across edge, fog, and cloud environments. It extends the SAREF4SYST ontology with concepts for AI pipelines, executable jobs, resources, deployment constraints, and communication links, creating a unified semantic model for both AI workflows and heterogeneous infrastructures. Evaluation through smart‑grid energy service scenarios and competency questions demonstrates successful deployment, reasoning, and workload adaptation, achieving 90‑100% deployment success and sub‑80 ms orchestration times.

By Viorica Rozina Chifu, Tudor Cioara, Vasile Ofrim, Liana Toderean, Ionut Anghel, Laura Daniele, Cornelis Bouter
arXiv AI
Jun 29

Agent-as-a-Router: Agentic Model Routing for Coding Tasks

arXiv:2606. 22902v3 Announce Type: replace Abstract: Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all.

By Pengfei Zhou, Zhiwei Tang, Yixing Ma, Jiasheng Tang, Yizeng Han, Zhenglin Wan, Fanqing Meng, Wei Wang, Bohan Zhuang, Wangbo Zhao, Yang You