The paper investigates when reallocating a fixed test‑time budget toward harder instances improves solution quality for neural combinatorial optimization solvers. Through pre‑registered experiments on three solvers and two hard‑workload constructions for the traveling salesman problem, it finds that the key deciding factor is the variation in instance difficulty within a workload, not the average difficulty. A budget‑aware policy that first spends part of the budget to gauge instance difficulty recovers most of the potential improvement, though not all, when the cost of this information is included.
By Jinhyung Bae
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.
By Barada Sahu (Cabal AI), Shivesh Pandey (Para AI)
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.
By Yunpeng Xu, Kun Zheng
arXiv:2609.26272v1 Announce Type: new
Abstract: Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to...
By Jian Xu
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
By Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin