arXiv:2609. 10954v1 Announce Type: new Abstract: Continual world models must decide whether new data justify changing the model.
By Anqi Peter Li, Kaden Kim
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.
By Barada Sahu (Cabal AI), Shivesh Pandey (Para AI)
The paper investigates how to handle a large‑language‑model’s weak verifier that accepts a response, followed by a second call that may resample or reroute. It introduces three decision gates—recoverable stopping debt, two‑sided FIT action support, and held‑out value from an outcome‑blind selector—to determine action selection, which is treated as an identification problem. Experiments on MBPP+, LiveCodeBench, and BigCodeBench show that while stopping debt exists, the current evidence does not identify when to resample versus reroute.
By Teng-Ruei Chen
The study investigates how inference‑time interventions and weight consolidation affect open‑ended generation in an online bin‑packing task. By iteratively generating, verifying, selecting, and consolidating with LoRA, the model’s outputs shift toward higher value, reducing excess by 1.7 points and outperforming random consolidation by 3.1 points. Across three independent runs, the mean performance remained consistent, and the best candidates converged to the classic heuristic’s level without exceeding it, while consolidation also lowered the proportion of better‑than‑classic candidates but increased their absolute number.
By Roberto I. Ono Filho
The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.
By Michel A. Youssef
arXiv:2609.07162v1 Announce Type: new
Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
By Xin Xu