arXiv:2604. 22027v2 Announce Type: replace-cross Abstract: One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed.
By Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez, Eric Todd, David Bau, Michael L. Littman, Stephen H. Bach, Ellie Pavlick
The study turns the speculative field of astrolinguistics into an experiment by using two large language models with deliberately incompatible constructed lexicons as informants. A scripted orchestrator translates between the two category systems, and a protocol combining cross‑situational elimination, predictive probes, active scene selection, and a stricter recovery round successfully prevents the ‘kangaroo effect’—the silent attachment of a word to the wrong referent—in over 400 simulated and live runs. When informant noise is introduced, the protocol remains robust up to 2% per‑word noise and largely abstains rather than errs at higher noise levels, while a generate‑and‑test loop allows recovery of words outside the scripted hypothesis space, achieving full coverage as the rule‑proposing LLM’s capability increases.
whyItMatters":"The protocol demonstrates that experimental astrolinguistics can reliably avoid mistranslations and recover unknown terms, showing that correctness is governed by the protocol while coverage depends on the instruments used."
By Francesco Cordella, Mauro Cappelli
The paper investigates how language models can covertly encode a hidden trait—termed subliminal learning—through seemingly unrelated outputs. By systematically measuring output co‑variation, fixed output‑vector alignment, hidden‑state readability, and causal control across a range of model sizes and prompting protocols, the authors find that fixed geometry and observational readability do not reliably predict behavior, while causal timing and multi‑token measurements reveal stronger, concept‑wide effects. These distinct properties highlight that token‑level explanations are insufficient to pinpoint the mechanism behind training‑time trait transfer.
By Barath Velmurugan
arXiv:2506.17871v4 Announce Type: replace-cross
Abstract: Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this cons...
By Chenghao Yang, Sida Li, Ari Holtzman
arXiv:2608. 04021v1 Announce Type: cross Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits.
By Han-yu Wang
arXiv:2608. 14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior.
By Jinglei Ren, Yuyue Wang