arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
By Aimen Boukhari
arXiv:2510. 20535v2 Announce Type: replace-cross Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs.
By Hippolyte Pilchen, Edouard Grave, Patrick P\'erez
The study investigates which neurons in a frozen BERT-base-uncased encoder support AI‑text detection using the RAID benchmark across six generators. By applying an L1‑to‑L2 sparse‑probing protocol to all 9,216 CLS hidden‑state dimensions, the authors identify a stable set of fewer than 1% of neurons per generator that largely preserves detection accuracy. Bidirectional activation patching confirms the causal relevance of this set, while mean‑ablating the neurons shows the signal is redundantly distributed, and cross‑generator analysis reveals a bipartite structure with instruction‑tuned generators concentrating more stable neurons in the final layer.
"whyItMatters":"The findings demonstrate that a small, stable subset of BERT neurons can reliably support AI‑text detection across diverse generators, enabling efficient detector design without re‑identifying neurons for each new generator."
By Pawe{\l} Blicharz, Mi{\l}osz Grunwald
arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.
By Albert Tseng, Christopher De Sa
arXiv:2609.35232v2 Announce Type: replace-cross
Abstract: Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pr...
By Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
arXiv:2607. 00004v1 Announce Type: cross Abstract: While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the aging BERT-base baseline in learned sparse retrieval (LSR).
By Zhichao Geng, Yang Yang