arXiv:2604. 07472v2 Announce Type: replace Abstract: Serving large language model (LLM) inference in cloud environments requires jointly optimizing model selection, GPU provisioning, parallelism configuration, and workload routing under latency, accuracy, memory, and budget constraints.
By Jiaming Cheng, Duong Tung Nguyen
arXiv:2607. 06066v1 Announce Type: new Abstract: The Vehicle Routing Problem (VRP) and its variants represent some of the most practically consequential optimization challenges in modern logistics and urban mobility.
By Manish Kolachalam, Rani Malhotra
arXiv:2607. 03694v1 Announce Type: new Abstract: Large-scale Capacitated Vehicle Routing Problems (CVRPs) are commonly solved by partitioning customers into smaller routing problems that can be optimized independently.
By Oguzhan Karaahmetoglu, Hyong Kim
arXiv:2607. 13080v1 Announce Type: cross Abstract: Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost -- and on-premise quantized open-weights models, which promise low-marginal-cost scaling and data sovereignty at some loss of reasoning fidelity.
By Sheng-Wei Peng, Yi-Hsun Lin, Yi-Pei Lee
arXiv:2608. 11361v1 Announce Type: new Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis.
By Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu
arXiv:2606. 13241v1 Announce Type: new Abstract: Defining query difficulty is one of the hardest problems in deployment engineering.
By Francesco Massa, Marco Cristofanilli