arXiv AI

FFNet: MetaMixer-based Efficient Convolutional Mixer Design

arXiv:2406. 02021v3 Announce Type: replace-cross Abstract: Transformer, composed of self-attention and Feed-Forward Network, has revolutionized the landscape of network design across various vision tasks.

arXiv Machine Learning
Jun 3

Dynamic Short Convolutions Improve Transformers

arXiv:2606. 03825v1 Announce Type: new Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization.

By Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim
arXiv Computation and Language
Aug 25

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

The paper investigates whether deep transformer layers require context from the residual stream to compute value vectors. It finds that allowing deeper layers to use a context‑free value vector—preserving original token information—significantly improves performance, and adding context afterward yields little extra benefit. The authors introduce Bank of Values (BoV), a lookup table of token‑specific value vectors for the last third of layers, which reduces compute and memory while matching or surpassing prior methods on large models.

By Muyu He, Yuchen Liu, Qingya Huang, Li Zhang
arXiv Machine Learning
Jul 14

Controllably Efficient Language Models

arXiv:2511. 05313v2 Announce Type: replace Abstract: The substantial inference costs of attention in transformers motivated the development of efficient sequence mixers: namely sparse and sliding window attention, convolutions and linear attention.

By Jatin Prakash, Aahlad Puli, Rajesh Ranganath
arXiv AI
Sep 16

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

QueryFormer is a unified architecture designed for post‑click conversion rate prediction, addressing both feature interactions and sequential user behaviors. It introduces a stackable field–sequence block that generates query tokens via cross‑attention and packs sequence queries into shared‑parameter attention, improving efficiency and accuracy. The model won first place in the KDD Cup 2026 Tencent UniRec Challenge Industrial Track with an AUC of 0.83254, and scaling studies show that increasing view width slightly boosts validation AUC while maintaining low latency.

By Yuanzhe Zhou, Zhaoyang Zeng
arXiv Computation and Language
6d ago

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

The paper proposes a new architecture for masked language modeling that replaces the Transformer attention mechanism with a stack of low‑rank bottleneck autoencoders. Each autoencoder mixes information locally, across the full sequence, and across attention heads, compressing and reconstructing inputs without training‑dependent width. An iterative refinement process at masked positions pulls embeddings toward a weighted neighbor average and then projects them back onto the learned manifold, achieving comparable performance to BERT with roughly 1.9× fewer FLOPs and matching BERT on rare‑token performance through a frequency‑aware training schedule.

By Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
arXiv AI
2d ago

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation. "whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."

By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen
arXiv AI
Jul 28

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

arXiv:2607. 22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes.

By Xin Yang, Yemin Wang, Mingda Liu, Letian Li, Shuaishuai Cao, Zhengxiao He, Ryan Dong