arXiv:2608. 16947v1 Announce Type: cross Abstract: Huang, Lou, and Xiao introduced Dynamic Mixture-of-Experts Serving and gave an O(sqrt(log k))-competitive randomized algorithm for its integral primal problem, where k is the number of replica GPUs beyond the mandatory copy of each expert.
By Ian D'Ambrosio (Nth Research Collective)
arXiv:2608.12134v2 Announce Type: replace-cross
Abstract: We study nonnegative submodular maximization on $n$ elements subject to a general matroid of rank $k$, when the offline algorithm is given an...
By Vaneet Aggarwal
arXiv:2603. 00910v2 Announce Type: replace-cross Abstract: Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas others are nearly redundant.
By Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra, Ankur Mali
The paper proposes a method for routing requests to a fixed pool of quantized Mixture-of-Experts (MoE) instances, aiming to maximize throughput while respecting a quality‑degradation budget. It introduces Fragility‑Weighted Perplexity (FWP) as a request‑specific risk metric derived from prompt tokens, and uses a window‑level linear program to compute a reduced‑reward score that aligns with the LP optimum. Experiments on Qwen prompts show that FWP‑based allocation improves throughput by 2.5% over request‑agnostic mixing and static configurations.
By Zhenghong Huang, Hongfan Wu, Jiheng Zhang
arXiv:2608. 07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply.
By Simone Mainardi, Kaushal Bansal, Prabhat Singh
arXiv:2608.29097v1 Announce Type: cross
Abstract: This paper studies the problem of proportionally fair clustering, where the goal is to select $k$ ``centers'' from a metric space that fairly represe...
By Benjamin Cookson, Eva Deltl, Yeeseok Oh
arXiv:2608. 18130v1 Announce Type: cross Abstract: We study online bipartite matching with reusable server capacity and non-stationary rewards.
By Xi Chen, Shixin Wang, Bingkun Zhou, Yuan Zhou
arXiv:2608. 07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast.
By Mojtaba Eslami
arXiv:2606. 00835v1 Announce Type: new Abstract: Network routers that enforce Quality-of-Service (QoS) guarantees must decide, at every clock cycle, which expiring packet of information to transmit, even when the value of the packet is unknown until it is processed.
By Gianmarco Genalti, Achraf Azize, Vianney Perchet
The paper presents a block‑wise differentiable Sinkhorn attention mechanism designed for long‑context balanced entropic optimal transport on TPU hardware. By stopping a $T$‑step Sinkhorn solve and unrolling a short refinement tail, the authors derive an exact surrogate gradient that achieves efficient block‑wise cost and memory usage. Experimental results on synthetic masked problems and a Pfam protein‑family screen demonstrate high numerical accuracy and sustained throughput on TPU v6e‑8, with notable improvements in reconstruction and sparse cross‑entropy metrics.
By Dylan Forde
arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.
By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
arXiv:2608. 12134v1 Announce Type: cross Abstract: We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value oracle.
By Vaneet Aggarwal