One mechanism for many mental spaces: a shared router over a value slot in language models
arXiv:2607. 10248v1 Announce Type: cross Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical.
arXiv:2607. 10248v1 Announce Type: cross Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical.
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
The paper investigates how language models decide between contextual information and their internal memory when the two conflict. By estimating "authority directions" from agreement prompts and swapping these directions between matched prompts, the authors show that such interventions can reproduce 30–68% of the shift in source choice across Qwen, Llama, and OLMo models. Cross‑task experiments reveal that authority directions learned on one task transfer only modestly (≈9%) to another, indicating that authority computations are largely task‑specific.
A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.
arXiv:2607. 11020v1 Announce Type: cross Abstract: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights.
arXiv:2607. 11945v1 Announce Type: cross Abstract: Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world.
The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.
The paper investigates how large language models (LLMs) such as Qwen, Llama, and Gemma use internal knowledge when answering questions. By performing layer‑wise interventions on the hidden state after the question, the authors compare how different request directions (pair‑conditioned vs. global) and answer types (noun, adjective, code) influence the model’s routing of information. The study finds that the influence of request direction varies across models and layers, with some models showing a sustained routing effect while others do not, highlighting distinct patterns of early readability, causal steering, and later content dependence.
arXiv:2606. 15733v1 Announce Type: cross Abstract: Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged.
arXiv:2601. 06599v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations.
arXiv:2602.22453v4 Announce Type: replace Abstract: Retrieval heads, a subset of attention heads in Transformers, were studied in English, showing its crucial role in retrieving information from the...
arXiv:2606. 15405v1 Announce Type: cross Abstract: Long-term memory is essential for conversational agents to remain coherent across extended dialogues, follow through on commitments made many sessions earlier, and adapt their behaviour to each user.