arXiv Machine Learning By He Zhang

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

Read the original on arXiv Machine Learning →

TriPLU is a Trilinear Product Linear Unit that replaces the gated feed‑forward branch in tiny decoder‑only language models with a degree‑3 product‑only branch that multiplies three projected streams coordinate‑wise. In a character‑level TinyStories 1M‑byte prefix study, TriPLU achieves a mean best validation loss of 1.0637, outperforming closely matched SwiGLU (1.1017), a degree‑4 product control (1.0780), and a degree‑2 control (1.1026). In train‑only Byte‑BPE experiments, TriPLU also lowers validation and held‑out bits per byte on TinyStories and WikiText‑2 raw under low‑learning‑rate settings, with PMI‑slice evidence indicating gains on seen middle‑ and high‑PMI adjacent‑token pairs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 14

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

The paper investigates whether small models distilled from larger ones behave similarly when using byte versus token tokenization. It introduces two methods—Marginalize‑It (approximate) and End‑Of‑Token (exact)—to convert token logits to byte logits, and conducts a large‑scale study on decoder‑only dense transformers ranging from 1 billion to 1 trillion bytes of data. Results show that while token‑based models excel early, byte‑based models eventually surpass them with more compute, achieving higher performance ceilings, greater data efficiency, and lower logit storage costs.

By Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer
arXiv Machine Learning
1d ago

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

The paper investigates how exporting ternary language models (BitNet, Falcon‑E, BitCPM) through a bf16 cast step can introduce significant discrepancies between the fine‑tuned latent weights and the deployed ternary codes. In three lab pipelines, the authors find that fp32 quantization of shipped latents disagrees with the deployed codes on up to 1.77% of codes, and that the export step can drastically reduce strict accuracy on GSM8K (e.g., from 58.79% to 0.78% for Falcon‑E‑1B‑Base). They propose two compatibility remedies—directly writing the training quantizer’s codes or adjusting bf16 inputs—to meet a 4‑point strict‑accuracy non‑inferiority criterion across all models.

By Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness)
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
arXiv Computation and Language
Aug 31

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

The paper demonstrates that a byte‑level BPE tokenizer can be sliced to create multiple vocabulary sizes from a single trained model, preserving exact logits while reducing deployed weights by 66%. Experiments on 30 models show that while sliced models match the full model numerically, they underperform fixed‑cap specialists by a few percentage points in bits‑per‑byte. Multi‑cap training improves robustness to typographical noise, suggesting benefits from training across multiple granularities rather than from control tokens alone.

By Christos Koutsiaris