arXiv AI By Narcis Marincat

What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies

Read the original on arXiv AI →

arXiv:2608. 20054v1 Announce Type: new Abstract: Multi-module systems often expose every module to the full input.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 7

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.

By Ayushi Agarwal
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv AI
2d ago

First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers

The paper introduces Language Identity Head Ablation (LIHA), a causal method that zeroes individual attention heads in transformer models to measure language switch rates across multilingual prompts. Applying LIHA to GPT‑2 reveals a small set of first‑token broadcaster heads—most notably L6H1—that persistently attend to the initial prompt token and propagate language signals throughout generation, with compensatory head recruitment occurring hierarchically in higher layers. A controlled comparison between Qwen2.5‑1.5B‑Base and Qwen2.5‑1.5B‑Instruct shows that instruction tuning concentrates language‑identity influence in early layers, while experiments with Chinese and Russian confirm script‑specific first‑token broadcasting at layer 0.

By Arjun Pillai, Christian Hoang, Anjelo Jann Laroza