arXiv:2605. 20982v2 Announce Type: replace-cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement, adaptive expert relayout, hierarchical collectives, and EP-aware topology.
By Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein
The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.
By Yunpeng Xu, Kun Zheng
The study examines the composition of a random sample from the Model Context Protocol (MCP) registry, revealing that only 48.8% of the 400 sampled npm/stdio servers successfully complete an initialization handshake, compared to 66.7% for a hand‑curated frame. Among the servers that run, there are no fatal JSON Schema violations across 2,766 advertised tools, but optional safety annotations vary widely, with a 58.8% omission rate in the random draw versus 41.5% in the curated set. The authors also compare MCP tool descriptions to two benchmark corpora, finding minimal near‑duplication in real MCP tools (2.8%) and significant repetition in synthetic datasets (up to 85.6%).
By Haseeb Mohammed Afsar
The paper introduces the Compute-Value Audit (CVA), a sequential framework that evaluates whether extra sampling during test‑time scaling for video world models actually yields a net benefit after accounting for the compute cost of generation and verification. On 192 Physics‑IQ scenes, increasing the sample pool from 4 to 16 candidates improves oracle quality by +9.23 IQ, yet common metrics such as Flow, Cycle, and VideoReward fail to reliably recover this headroom, and adaptive‑depth policies recover only 42‑69% of the potential gain. Only a few specific interventions—anchor‑explorer in a sparse PRM800K setting, MMLU‑Pro exposing a predictive‑state gap, and a privileged paired‑future upper bound—successfully pass all CVA stages, indicating that sampling headroom is valuable only when it can be converted into a reliable decision that survives the full compute charge.
By Yuhua Jiang, Junjie Lu, Feifei Gao
The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.
By Tanay Chowdhury, Saeideh Shahrokh Esfahani
arXiv:2607. 17205v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) of open-weight LLMs on expert agent trajectories has emerged as a prominent approach to building capable code agents without reliance on proprietary models.
By Yunze Han