arXiv AI By Megi Dervishi, Mathurin Videau, Yann LeCun

Separating Representation from Reconstruction Enables Scalable Text Encoders

Read the original on arXiv AI →

arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

The study investigates which neurons in a frozen BERT-base-uncased encoder support AI‑text detection using the RAID benchmark across six generators. By applying an L1‑to‑L2 sparse‑probing protocol to all 9,216 CLS hidden‑state dimensions, the authors identify a stable set of fewer than 1% of neurons per generator that largely preserves detection accuracy. Bidirectional activation patching confirms the causal relevance of this set, while mean‑ablating the neurons shows the signal is redundantly distributed, and cross‑generator analysis reveals a bipartite structure with instruction‑tuned generators concentrating more stable neurons in the final layer. "whyItMatters":"The findings demonstrate that a small, stable subset of BERT neurons can reliably support AI‑text detection across diverse generators, enabling efficient detector design without re‑identifying neurons for each new generator."

By Pawe{\l} Blicharz, Mi{\l}osz Grunwald
arXiv AI
Jun 4

L$^3$: Large Lookup Layers

arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.

By Albert Tseng, Christopher De Sa