arXiv AI

What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies

arXiv AI
Sep 17

What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

The study investigates whether restricting a module’s access to information—through evidence masking—enhances a system’s ability to learn compositional tasks. In a preregistered experiment with sixty‑four‑cell systems built on a frozen language‑model backbone, researchers varied evidence masking, ownership markers, and filler replacement across multiple initialization clusters and data orders. Results showed that when markers were available, masking significantly improved accuracy on held‑out two‑ and three‑operation compositions, with all tested pairs meeting performance thresholds and the preregistered behavioral criterion satisfied. The study also explored packet interventions and found predicted intermediate‑value changes, though mediation was not conclusively established. "whyItMatters":"The findings demonstrate a substantial performance benefit from evidence masking in compositional generalization tasks, offering a promising direction for designing more effective learning systems."

By Narcis Marincat
arXiv AI
Sep 12

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.

By Narcis Marincat
Hugging Face Trending Papers
Sep 10

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that most cross‑initialization pairs fail to interoperate, with only one pair achieving full bidirectional compatibility. Further experiments reveal that reinitializing only the packet reader, writer, and mouth dramatically improves accuracy, while inherited interfaces never outperform fresh ones by the preregistered margin.

arXiv Machine Learning
Aug 7

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.

By Ayushi Agarwal
arXiv AI
Sep 24

ProteinJEPA: Latent prediction improves protein language model pretraining

ProteinJEPA introduces a joint‑embedding predictive architecture that supplements masked language modeling (MLM) with a cosine loss to predict latent representations of a teacher model. On 19 protein tasks, MLM+JEPA outperforms compute‑matched and step‑matched MLM‑only training across 78 and 76 of 114 comparisons, achieving notable gains on structure‑ and homology‑sensitive tasks such as SCOPe‑40 retrieval and remote homology. Ablation studies show the cosine loss is superior to mean squared error and that latent prediction complements rather than replaces MLM.

By Dan Ofer, Dafna Shahaf, Michal Linial
arXiv AI
Sep 16

TAME: Token Attribution and Masking for Emergent misalignment

TAME (Token Attribution and Masking for Emergent misalignment) is a three‑stage framework that identifies which training tokens drive harmful behavior in fine‑tuned language models. It first scores tokens by how much fine‑tuning increases their likelihood, then characterizes patterns among high‑attribution tokens, and finally validates them by masking during training. Experiments on Llama and Qwen show that masking the top‑attribution tokens reduces emergent misalignment by 23‑ to 36‑fold, while random masking has no effect.

By Md Rayhanul Masud, Md Rizwan Parvez
arXiv Machine Learning
Sep 18

Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds

The paper investigates how language models can covertly encode a hidden trait—termed subliminal learning—through seemingly unrelated outputs. By systematically measuring output co‑variation, fixed output‑vector alignment, hidden‑state readability, and causal control across a range of model sizes and prompting protocols, the authors find that fixed geometry and observational readability do not reliably predict behavior, while causal timing and multi‑token measurements reveal stronger, concept‑wide effects. These distinct properties highlight that token‑level explanations are insufficient to pinpoint the mechanism behind training‑time trait transfer.

By Barath Velmurugan