arXiv:2608. 14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior.
By Jinglei Ren, Yuyue Wang
arXiv:2608. 12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.
By Arnav Srivastav
The paper investigates neural text degeneration by measuring the fixed‑point structure of short‑window argmax maps across 17 pretrained models, using 96 random two‑token starts without prompts. It finds a stable four‑way classification that varies across model families and scales, with some models funneling to a single endpoint token while others do not, and shows that this behavior is not solely determined by training data or corpus frequency. The study demonstrates that repetition phenomena are not uniformly explained by either training data or network architecture alone, highlighting the complexity of neural text generation dynamics.
By Nicol\'as Vera Z\'u\~niga
Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can...
arXiv:2606. 24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition.
By Jessica Chudnovsky, Joshua Kazdan, Noam Levi, Rylan Schaeffer, Yegor Denisov-Blanch, Bo He, Mehmet Donmez, Sanmi Koyejo, David Donoho
The paper studies structural priming in language model production by conducting controlled sentence‑completion experiments on dative constructions. Results show that language models exhibit priming effects, especially when sentences are semantically coherent, with stronger relative increases for double‑object datives and larger absolute increases for prepositional‑object datives. The study also finds that primed completions involve more lexico‑semantic repetition, indicating that priming operates across syntactic, lexical, and semantic levels.
By Giulia Pucci, Ruizhe Li, Arabella Sinclair
The paper introduces MWE‑ECL, a bilingual diagnostic framework that tests whether distant discourse anchors can override local lexical priors in multi‑word expression interpretation. It evaluates models on a 0‑128K context grid, finding that while retrieval of anchors is near perfect, the ability to change locally preferred readings varies, especially when the model’s default conflicts with the anchor. The study shows that explicit recoverability does not always translate into behavioral influence, with gaps differing across models and languages.
By Wei He, Aline Villavicencio, Rodrigo Wilkens, Zhenyun Deng
arXiv:2608. 04330v1 Announce Type: cross Abstract: Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens with little change.
By Mike Vegeto
arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.
By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.
By Arman Nik Khah
Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens with little change. We turn this observation into prefix-removal probing and introduce Right Reset (RR), which measures preservation of the right-hand hidden-state trajectory.
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
By Zhenyan Lu, He Wang, Xiaohui Huang