LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolatio...
The paper evaluates a learned request‑routing policy for disaggregated large‑language‑model serving, where compute‑heavy prefill and memory‑heavy decode stages run on separate GPU pools. Using a discrete‑event simulator and real NVIDIA A40 GPUs, the calibrated router—leveraging prompt length, predicted output length, KV‑cache pressure, and SLO class—outperforms round‑robin, least‑loaded, and length‑based heuristics, achieving the highest mean goodput (0.864) and lowest variance across three mixed, bursty arrival traces. Hardware calibration proves critical, providing a 4.5‑point goodput boost and roughly 40 % of the tail‑latency advantage, and the learned router can match round‑robin performance with one fewer GPU in certain scenarios.
By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
The paper re‑examines the impact of removing the encoder from Action Chunking Transformers (ACT), a model used for robot manipulation learning. Contrary to the original claim that encoder removal drops success rates from 35% to 2%, the authors find no such dramatic effect in their re‑runs, though minor variations remain uncertain. They attribute discrepancies to training length and checkpoint selection, and note that the encoder’s latent variable offers little reconstruction benefit on the tested benchmark, while its removal speeds up training.
By Bo Kang
The paper introduces MedPCFM‑TED, a one‑step distillation framework that uses teacher‑guided endpoint supervision and geometric matching losses to generate cranial implants from point clouds. It outperforms existing one‑step methods on the SkullBreak benchmark, remains competitive on SkullFix, and achieves a generation time of about 0.04 s per sample. The approach demonstrates that rapid, high‑quality implant generation is possible without multiple neural evaluations during inference.
By Kamil Kwarciak, Marek Wodzinski
The paper introduces LCAP, a method for adapting photonic neural networks to real hardware by learning a shared correction from a population of chips and then personalizing each chip using only 32 fixed output probes. LCAP decomposes adaptation into a transferable population correction and a probe‑inferred latent personalization, allowing feed‑forward calibration without device‑specific optimization. Experiments on a simulated three‑layer 64‑mode MZI network show accuracy improvements from 80.4% to 93.4% and significant gains on unseen chips.
By Tianyu Gao, Guantian Zheng
The study investigates how pruning affects large language models (LLMs) used for smart‑home tool calling. Researchers examined four LLMs—dense Transformer, dense hybrid, and mixture‑of‑experts (MoE) architectures—using depth, width, hybrid, and expert pruning, followed by supervised fine‑tuning. They evaluated over 19,500 instances from three smart‑home datasets, analyzing not only overall accuracy but also degradation in action components (operation, device, argument, value) and task complexity, finding that dense models suffer sharp performance drops after a narrow safe pruning range, while MoE models tolerate more pruning; aggressive pruning also leads to over‑refusal and loss of grounded specificity.
By Congjing Zhang, Vashishtha Patil, Henning Lange, Usman Aleem
arXiv:2609.16637v1 Announce Type: cross
Abstract: Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larg...
By Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia
arXiv:2609.17026v1 Announce Type: new
Abstract: Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a da...
By Yunxiang Fu, Meng Lou, Zicheng Liao, Yizhou Yu
arXiv:2609.17509v1 Announce Type: cross
Abstract: Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing comput...
By Thanapat Trachu, Samuele Cornell, William Chen, Shinji Watanabe
arXiv:2609.17109v1 Announce Type: new
Abstract: A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefi...
By Dushyant Rajput
arXiv:2609.17391v1 Announce Type: new
Abstract: Model serving is one of the largest cost drivers in production recommender systems. Maximizing its throughput requires navigating a deeply layered hier...
By Qi Wu, Lohan Lemire, Kai Meng, Zhongmou Cai, Raphael Bargues, Petr Zhitnikov, Zeyuan Cao, Yao Wang, Shujun Bian, Wei Chen, Sean Sheng
arXiv:2609.16215v1 Announce Type: new
Abstract: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate s...
By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv:2609.17019v1 Announce Type: new
Abstract: While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference...
By Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang, Linna Zhou
arXiv:2609.16255v1 Announce Type: cross
Abstract: We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Ou...
By Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania
arXiv:2609.16450v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative...
By Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, Haoran You
arXiv:2607.09709v2 Announce Type: replace
Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study th...
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2609.16689v1 Announce Type: new
Abstract: Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained...
By Jinwoo Jeon, GyuYeop Do, Yubin Lim, Nam-Joon Kim, Hyun Gon Ryu, Hyuk-Jae Lee, Byung-Jun Lee
arXiv:2609.16874v1 Announce Type: new
Abstract: Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of th...
By Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool
arXiv:2507.01927v3 Announce Type: replace
Abstract: While CNNs and ViTs dominate vision architectures, all-MLP models offer a structurally simpler alternative whose patch-independent processing is na...
By Zhentan Zheng
The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.
By Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami