arXiv AI

Separating Representation from Reconstruction Enables Scalable Text Encoders

arXiv:2607. 04011v1 Announce Type: cross Abstract: While decoders have rapidly scaled, encoders have remained largely unchanged since BERT.

arXiv AI
6d ago

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

The study investigates which neurons in a frozen BERT-base-uncased encoder support AI‑text detection using the RAID benchmark across six generators. By applying an L1‑to‑L2 sparse‑probing protocol to all 9,216 CLS hidden‑state dimensions, the authors identify a stable set of fewer than 1% of neurons per generator that largely preserves detection accuracy. Bidirectional activation patching confirms the causal relevance of this set, while mean‑ablating the neurons shows the signal is redundantly distributed, and cross‑generator analysis reveals a bipartite structure with instruction‑tuned generators concentrating more stable neurons in the final layer. "whyItMatters":"The findings demonstrate that a small, stable subset of BERT neurons can reliably support AI‑text detection across diverse generators, enabling efficient detector design without re‑identifying neurons for each new generator."

By Pawe{\l} Blicharz, Mi{\l}osz Grunwald
arXiv AI
Jun 4

L$^3$: Large Lookup Layers

arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.

By Albert Tseng, Christopher De Sa
arXiv AI
Jun 10

PromptEmbedder: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting

arXiv:2605. 28066v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottlenecks in computational efficiency and cross-architecture transferability.

By Yu-Che Tsai, Kuan-Yu Chen, Yuan-Hao Chen, Yu-Han Chang, Ching-Yu Tsai, Yu-Hsiang Chuang, Shou-De Lin
arXiv Machine Learning
Sep 11

SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations

SG-Blend introduces a per‑layer adaptive activation that interpolates between a bias‑corrected, parametric Swish variant (SSwish) and GELU, using a learnable blend coefficient, sharpness, and zero‑centering bias. The method adds only three scalars per feed‑forward block and, on BERT‑style IMDB classification, matches peak accuracy while reducing seed‑to‑seed variance by 42 %. It also achieves the lowest validation perplexity on WikiText103 and generalizes to computer vision and other domains.

By Gaurav Sarkar, Syed Affan Daimi, Jay Gala, Subarna Tripathi
arXiv AI
Jun 16

Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality

arXiv:2505. 18227v4 Announce Type: replace-cross Abstract: In Transformer architectures, tokens\textemdash discrete units derived from raw data\textemdash are formed by segmenting inputs into fixed-length chunks.

By Zhenglun Kong, Yize Li, Fanhu Zeng, Lei Xin, Shvat Messica, Xue Lin, Pu Zhao, Manolis Kellis, Hao Tang, Marinka Zitnik