The paper investigates whether small models distilled from larger ones behave similarly when using byte versus token tokenization. It introduces two methods—Marginalize‑It (approximate) and End‑Of‑Token (exact)—to convert token logits to byte logits, and conducts a large‑scale study on decoder‑only dense transformers ranging from 1 billion to 1 trillion bytes of data. Results show that while token‑based models excel early, byte‑based models eventually surpass them with more compute, achieving higher performance ceilings, greater data efficiency, and lower logit storage costs.
By Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer
Orthrus is a hybrid autoregressive‑diffusion model that claims to perform lossless speculative decoding by generating multiple tokens in parallel while maintaining a frozen autoregressive backbone. In our independent reproduction, we found that under BF16 precision the model matches the exact autoregressive trajectory only about 45% of the time, whereas with FP32 precision it achieves perfect matching on all evaluated prompts. Despite trajectory mismatches at lower precision, Orthrus does not exhibit systematic degradation on downstream lm‑eval‑harness benchmarks.
By Ilya Koziev, Leonid Sinev, Ivan Oseledets
The paper argues that relying solely on zero‑shot task accuracy is insufficient for evaluating quantized large language models (LLMs) because accuracy ignores changes in the full predictive distribution. It proposes a distribution‑sensitive framework that measures fidelity loss by computing statistical distances—such as Jensen‑Shannon Divergence and Total Variation Distance—between the full‑vocabulary output distributions of a full‑precision BF16 reference and its quantized counterparts. Experiments across five foundation architectures and four reasoning benchmarks show that these divergence metrics increase with stronger quantization, revealing distributional drift that top‑1 accuracy fails to capture, and suggest that mixed‑precision Q4_K schemes can offer lower divergence than uniform Q4_0 at comparable memory usage.
By Shahzeb Qamar, Lorenz Sparrenberg, Christian Bauckhage, Baha Rababah, Carson Leung, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Rafet Sifa
The paper demonstrates that layer dropout, also known as stochastic depth, can be effectively used in state‑of‑the‑art large language model (LLM) training. By optimizing the layer distribution, schedule, and optimizer settings, the authors show that layer dropout can reduce training loss while saving up to 25 % of training FLOPs. Additionally, layer dropout enables post‑training optimizations such as early exit and self‑speculative decoding, achieving up to 1.5× inference speedup with negligible accuracy loss across models ranging from 271 M to 8.2 B parameters and datasets up to 160 B tokens.
By Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal, Abdulrahman Mahmoud, Shane Bergsma, Joel Hestness
Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by generating multiple tokens in parallel while using a frozen autoregressive backbon...
arXiv:2607. 14427v1 Announce Type: new Abstract: A depth-recurrent transformer applies a weight-tied core a variable number of times, and prior work has shown that training with a randomized recursion count yields one checkpoint usable across a range of inference depths.
By Joe Logan
Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.
By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv:2607. 09999v1 Announce Type: cross Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved.
By Renuka Oladri, Mohan Vamsi Varadaraju Priya, Jerry Wu
arXiv:2605. 09825v4 Announce Type: replace-cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable?
By Musa Cim, Sarthak Arora, Poovaiah Palangappa, Miro Hodak, Ravi Dwivedula, Meena Arunachalam, Mahmut Taylan Kandemir
arXiv:2608. 14563v1 Announce Type: cross Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.
By Rivaan Patil, Simon Dennis, Hao Guo, Kevin Shabahang
arXiv:2506. 21833v2 Announce Type: replace Abstract: Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing.
By Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan
The paper introduces Speculative Probing, a method that repurposes the speculative‑decoding module of large language models for real‑time classification tasks. By appending a trained soft prompt to the target sequence, the approach leverages the already‑cached KV store during inference, adding negligible overhead while achieving higher accuracy than traditional hidden‑state probes. Experiments on four classification tasks across multiple models show that these lightweight probes outperform zero‑shot GPT‑5.4‑mini and rival or surpass specialized 8B safety classifiers without running a full LLM.
By Collin Zhang, Tingwei Zhang, Vitaly Shmatikov