arXiv Machine Learning By Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness)

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

Read the original on arXiv Machine Learning →

The paper investigates how exporting ternary language models (BitNet, Falcon‑E, BitCPM) through a bf16 cast step can introduce significant discrepancies between the fine‑tuned latent weights and the deployed ternary codes. In three lab pipelines, the authors find that fp32 quantization of shipped latents disagrees with the deployed codes on up to 1.77% of codes, and that the export step can drastically reduce strict accuracy on GSM8K (e.g., from 58.79% to 0.78% for Falcon‑E‑1B‑Base). They propose two compatibility remedies—directly writing the training quantizer’s codes or adjusting bf16 inputs—to meet a 4‑point strict‑accuracy non‑inferiority criterion across all models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 30

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

arXiv:2607. 26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision.

By Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju Durai Raju
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv AI
Sep 3

Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

The paper reports a post‑training ternarization of the 4‑billion‑parameter Qwen model, achieving an effective 1.641‑bit representation for 81.62 % of its weights while keeping activations at 16‑bit precision. Accuracy drops from 64.5 % to 54.7 % across ten capability tests, with uneven degradation (e.g., BoolQ 84.6 % of teacher performance, ARC‑Challenge 43.8 %). After packing the ternary planes, the model size shrinks from 8.29 GiB to 3.96 GiB with negligible change in perplexity, though inference speed is not improved.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv Machine Learning
6d ago

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...

By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
arXiv Machine Learning
Aug 24

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

TriPLU is a Trilinear Product Linear Unit that replaces the gated feed‑forward branch in tiny decoder‑only language models with a degree‑3 product‑only branch that multiplies three projected streams coordinate‑wise. In a character‑level TinyStories 1M‑byte prefix study, TriPLU achieves a mean best validation loss of 1.0637, outperforming closely matched SwiGLU (1.1017), a degree‑4 product control (1.0780), and a degree‑2 control (1.1026). In train‑only Byte‑BPE experiments, TriPLU also lowers validation and held‑out bits per byte on TinyStories and WikiText‑2 raw under low‑learning‑rate settings, with PMI‑slice evidence indicating gains on seen middle‑ and high‑PMI adjacent‑token pairs.

By He Zhang