Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
arXiv:2609. 10954v1 Announce Type: new Abstract: Continual world models must decide whether new data justify changing the model.
arXiv:2609. 10954v1 Announce Type: new Abstract: Continual world models must decide whether new data justify changing the model.
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.
The paper investigates how to handle a large‑language‑model’s weak verifier that accepts a response, followed by a second call that may resample or reroute. It introduces three decision gates—recoverable stopping debt, two‑sided FIT action support, and held‑out value from an outcome‑blind selector—to determine action selection, which is treated as an identification problem. Experiments on MBPP+, LiveCodeBench, and BigCodeBench show that while stopping debt exists, the current evidence does not identify when to resample versus reroute.
The study investigates how inference‑time interventions and weight consolidation affect open‑ended generation in an online bin‑packing task. By iteratively generating, verifying, selecting, and consolidating with LoRA, the model’s outputs shift toward higher value, reducing excess by 1.7 points and outperforming random consolidation by 3.1 points. Across three independent runs, the mean performance remained consistent, and the best candidates converged to the classic heuristic’s level without exceeding it, while consolidation also lowered the proportion of better‑than‑classic candidates but increased their absolute number.
The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.
arXiv:2609.07162v1 Announce Type: new Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
arXiv:2606. 15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down.
arXiv:2609.01108v1 Announce Type: new Abstract: TRACE (Math & Lienhart, arXiv:2602.01135) reads causal graphs over event types out of a pretrained autoregressive sequence model by thresholding a per-...
The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
arXiv:2609. 17464v1 Announce Type: cross Abstract: Multi-agent systems split a task across a tree of agents and justify the split with folklore: smaller contexts, cleaner separation, parallelism.
The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.