Hugging Face Trending Papers

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that most cross‑initialization pairs fail to interoperate, with only one pair achieving full bidirectional compatibility. Further experiments reveal that reinitializing only the packet reader, writer, and mouth dramatically improves accuracy, while inherited interfaces never outperform fresh ones by the preregistered margin.

arXiv AI
Sep 12

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.

By Narcis Marincat
arXiv AI
Sep 17

What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

The study investigates whether restricting a module’s access to information—through evidence masking—enhances a system’s ability to learn compositional tasks. In a preregistered experiment with sixty‑four‑cell systems built on a frozen language‑model backbone, researchers varied evidence masking, ownership markers, and filler replacement across multiple initialization clusters and data orders. Results showed that when markers were available, masking significantly improved accuracy on held‑out two‑ and three‑operation compositions, with all tested pairs meeting performance thresholds and the preregistered behavioral criterion satisfied. The study also explored packet interventions and found predicted intermediate‑value changes, though mediation was not conclusively established. "whyItMatters":"The findings demonstrate a substantial performance benefit from evidence masking in compositional generalization tasks, offering a promising direction for designing more effective learning systems."

By Narcis Marincat
arXiv Machine Learning
Aug 17

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

arXiv:2608. 14465v1 Announce Type: cross Abstract: A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates.

By Ziyang Luo, Zhongyao Chu, Xinjie He, Youting Wang, Xukui Qin, Runxiong Wu, Yan-Syuan Chen