arXiv AI By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

Read the original on arXiv AI →

TinyCeNN-LM proposes a quality‑gated post‑training conversion framework that replaces attention in pretrained language models with CeNN‑inspired cellular‑recurrent layers. The method includes bounded local processing, compact recurrent memory, routing, fusion, and an accept‑or‑rollback validation step, with three specific implementations studied. Experiments on SmolLM2‑135M and Qwen3.5‑0.8B show that the conversion can accept certain layers while rejecting others based on representation fidelity and NLL thresholds, achieving minimal perplexity changes and modest downstream accuracy retention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

arXiv:2608.30320v1 Announce Type: new Abstract: We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and a...

By Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu
arXiv Machine Learning
Sep 18

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any-order decoding and strong performance with parallel decoding.

By Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv AI
Sep 4

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

The paper presents a 4‑bit quantization recipe, Minima: NVFP4 W4A4, that fully quantizes all linear layers—including the Gated DeltaNet (GDN) recurrent blocks—of the 27‑billion‑parameter Qwen3.8 LLM. Across a suite of benchmarks (perplexity, MMLU‑Pro, GSM8K, AIME'25, GPQA‑Diamond, LiveCodeBench, and RULER retrieval), the quantized model matches BF16 performance within seed noise while being 17.5 GiB in size and 14–19 % faster at prefill. The authors attribute this success to four mechanisms: block‑scaling of residuals, robust gate projections, the delta‑rule recurrence’s noise‑plateau behavior, and the per‑token quantization cost’s dilution over long contexts.

By Sergii Kozyrev, Davyd Maiboroda
Hugging Face Trending Papers
Sep 17

dQwen3.5: Hybrid-Attention Diffusion Language Models

The paper introduces dQwen3.5, a family of diffusion language models derived from the hybrid-attention architecture of Qwen3.5 at 0.8B, 2B, 4B, and 9B parameters. It demonstrates that adapting a hybrid AR backbone—combining attention and RNN layers—can be more efficient than full-attention models, reaching a target training loss in roughly half the tokens. Across scales, dQwen3.5 exhibits full-attention-like behavior in any‑order decoding and strong performance under parallel decoding.