The paper evaluates additive activation steering in chat and agent contexts, showing that the commonly used gain ratio (Δ_agent/Δ_chat) fails to reliably indicate potency and efficacy across multiple models and dose-response cells. By replacing the gain with a location metric, dEC50 (difference in EC50 between agent and chat), the authors demonstrate a more robust, two‑sided measure that consistently captures cross‑context shifts. The study also reports several refuted and unanswered claims, emphasizing that a single operating point cannot distinguish between displacement and gain effects.
By Lucas Pinto
The paper reports that agent evaluations often show a tool‑call rate of zero even when the model emits valid calls, because the interface censors the trajectory before downstream components see it. Experiments on BFCL v4 and tau‑bench demonstrate that swapping the serving adapter can change the observed call rate from 0.00 to 0.96/0.19 or from 0 to 636 calls, indicating that the interface—not the model—causes the discrepancy. A 98‑line preflight check is released to detect such silent failures, highlighting that tool‑call rates depend on the model‑interface stack rather than the model alone.
By Wenbo Wang
The study investigates which components of a neural network contribute to rapid generalization (grokking) and how stable that improvement remains during further training. By transferring internal attention and MLP weights along with token embeddings and readout, the authors achieve a 5.46‑percentage‑point boost in early accuracy and a 558‑step reduction in confirmation latency, while also demonstrating that freezing transferred representations largely prevents post‑grokking relapse. The work delineates clear component‑level differences between acceleration and stability, and identifies architectural limits where omitting donor embeddings leads to significant performance loss.
By Zeyu Jia
arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.
By Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
The paper introduces a schema‑adaptive action‑conditioned Joint‑Embedding Predictive Architecture (SAAC‑JEPA) for cross‑machine CNC transfer when only a subset of sensors overlap between source and target machines. Experiments show that pretraining does not improve source‑only forecasting, but a carefully selected action‑conditioned JEPA model achieves a zero‑shot RMSE of 0.546 on the target, outperforming persistence but falling short of certain baseline models. Ablation studies reveal that adding RevIN improves RMSE but harms calibration, and limited post‑lock adaptation can further reduce error.
By Ayoub Louaye Bouaziz, Matthieu Ostertag, Anton Demasles
arXiv:2607. 27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes.
By Christian Rosenthal