Interaction Scaling: Grounding the Third Axis of Test-Time Compute
arXiv:2607. 11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one.
There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal.
arXiv:2607. 11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one.
arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
The paper introduces Janus, a method for validating error patterns in language models by comparing error rates across predefined yes/no properties and using shuffled decoy labels to set significance thresholds. Janus requires that a pattern’s error difference surpasses the decoy-derived threshold and is replicated on held‑out data before reporting. Experiments on a controlled code‑finding task confirm several meaningful error patterns, while on MuSiQue and LongBench v2 Janus reports no confirmed patterns for the tested properties, contrasting with standard shuffling tests that sometimes confirm patterns.
arXiv:2607.01469v3 Announce Type: replace Abstract: Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice eval...
arXiv:2609.24194v1 Announce Type: new Abstract: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they c...
The paper investigates how providing execution traces to multimodal judges in agentic video‑generation systems can bias their verdicts. On a benchmark of 109 two‑event clips, traces that falsely report successful tool calls cause large‑language‑model judges to incorrectly accept 78–90 % of failures, while contradictory traces lead to 100 % rejection of correct clips. The effect persists even when judges are instructed to consider only the video frames, indicating that the vulnerability stems from the judges’ learned trust in tool logs rather than the visual content itself.
arXiv:2608. 06270v1 Announce Type: new Abstract: The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom.
arXiv:2607. 28576v1 Announce Type: cross Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought.
arXiv:2609.13308v1 Announce Type: cross Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...