arXiv Machine Learning By Lucas Pinto

Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy

Read the original on arXiv Machine Learning →

The paper evaluates additive activation steering in chat and agent contexts, showing that the commonly used gain ratio (Δ_agent/Δ_chat) fails to reliably indicate potency and efficacy across multiple models and dose-response cells. By replacing the gain with a location metric, dEC50 (difference in EC50 between agent and chat), the authors demonstrate a more robust, two‑sided measure that consistently captures cross‑context shifts. The study also reports several refuted and unanswered claims, emphasizing that a single operating point cannot distinguish between displacement and gain effects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 27

AgentDiff: Meaning-Bearing Rewrites Trigger Deeper Divergence than Presentation Changes in LLM Agents

The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”

By Liyun Zhang, Jiayi Guo
arXiv Machine Learning
Jun 2

Measuring the Symmetry--Data Exchange Rate

arXiv:2606. 01090v1 Announce Type: cross Abstract: Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds.

By Ahmed M. Adly
arXiv Machine Learning
Sep 22

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.

By Chencheng Zhu