arXiv:2609.36569v1 Announce Type: cross
Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...
By Yupeng Chang, Wenxuan Zhang, Yuan Wu
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
By Donna Vakalis
arXiv:2608. 12959v1 Announce Type: cross Abstract: Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades.
By Joyjeet Singh
arXiv:2606. 28876v3 Announce Type: replace-cross Abstract: Proposal.
By Junyi Zou, Avrova Donz
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2605.18163v2 Announce Type: replace
Abstract: Hallucination correction is not a one-direction problem. We show that intermediate layers are neither uniformly more truthful than final layers nor...
By Tej Sanibh Ranade
arXiv:2609.39934v1 Announce Type: cross
Abstract: Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable proba...
By Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan
arXiv:2609.24194v1 Announce Type: new
Abstract: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they c...
By Daein Weon, Dongho Kang
The paper introduces the Compute-Value Audit (CVA), a sequential framework that evaluates whether extra sampling during test‑time scaling for video world models actually yields a net benefit after accounting for the compute cost of generation and verification. On 192 Physics‑IQ scenes, increasing the sample pool from 4 to 16 candidates improves oracle quality by +9.23 IQ, yet common metrics such as Flow, Cycle, and VideoReward fail to reliably recover this headroom, and adaptive‑depth policies recover only 42‑69% of the potential gain. Only a few specific interventions—anchor‑explorer in a sparse PRM800K setting, MMLU‑Pro exposing a predictive‑state gap, and a privileged paired‑future upper bound—successfully pass all CVA stages, indicating that sampling headroom is valuable only when it can be converted into a reliable decision that survives the full compute charge.
By Yuhua Jiang, Junjie Lu, Feifei Gao
arXiv:2603.01209v3 Announce Type: replace
Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python v...
By Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen
The paper introduces an action‑class diagnostic framework for multi‑turn tool‑calling in large language model agents, breaking failures into action‑class miscalibration and action‑execution failure across a four‑class action space (TOOL_CALL, ASK, REFUSE, CONFIRM). It defines a self‑revealing upper bound (Acc GAR) to expose state‑grader masking of miscalibration and shows that miscalibration is a significant, previously hidden failure mode, especially for heavily tool‑trained families. The study demonstrates that calibration can be reshaped by context‑only perturbations, but the effects vary widely across models and perturbation mechanisms, underscoring the need for diagnostics beyond aggregate accuracy.
By Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin
arXiv:2609.08618v1 Announce Type: new
Abstract: Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this mis...
By Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang