arXiv:2607. 11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one.
By Bojie Li, Noah Shi
arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.
By Valentin No\"el
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.
By Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
By Esmail Gumaan
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
By Shikhar Shiromani, Leo Richter
The paper introduces Janus, a method for validating error patterns in language models by comparing error rates across predefined yes/no properties and using shuffled decoy labels to set significance thresholds. Janus requires that a pattern’s error difference surpasses the decoy-derived threshold and is replicated on held‑out data before reporting. Experiments on a controlled code‑finding task confirm several meaningful error patterns, while on MuSiQue and LongBench v2 Janus reports no confirmed patterns for the tested properties, contrasting with standard shuffling tests that sometimes confirm patterns.
By Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh