Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds
Read the original on arXiv Machine Learning →The paper investigates how language models can covertly encode a hidden trait—termed subliminal learning—through seemingly unrelated outputs. By systematically measuring output co‑variation, fixed output‑vector alignment, hidden‑state readability, and causal control across a range of model sizes and prompting protocols, the authors find that fixed geometry and observational readability do not reliably predict behavior, while causal timing and multi‑token measurements reveal stronger, concept‑wide effects. These distinct properties highlight that token‑level explanations are insufficient to pinpoint the mechanism behind training‑time trait transfer.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.