The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.
By Yunpeng Xu, Kun Zheng
The paper investigates whether sink-aware attention head selection remains valid after 4‑bit NF4 weight‑only post‑training quantization. Using Sink Topology Consistency metrics, it finds that global rank preservation stays high across Qwen2.5 and Llama‑3.2 models, yet top‑k head overlap drops to 61–79% and layer‑specific sink‑mass shifts can be substantial. The study also shows that cross‑domain calibration degrades more than within‑domain precision and that recalibration with a small number of samples can recover most of the stability, though full‑map stability may require updating more layers.
By Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou
The paper investigates how per-domain data composition during the mid‑training phase (between pre‑training and alignment) affects model performance. Experiments with Qwen3‑8B‑Base across five KOR‑Bench domains show that a moderate coverage band (10%‑40%) yields the best performance for each domain, and that alignment passes cannot fully close the gaps created by suboptimal mid‑training allocations. Additionally, zero coverage during mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.
The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.
By Zachary Speck, Asa Shepard
arXiv:2606. 12280v1 Announce Type: new Abstract: Post-training quantization lets large text-to-image diffusion transformers run on consumer GPUs, yet the hardware-specific trade-offs are seldom measured directly.
By Deep Gandhi, Ali Asaria, Tony Salomone
arXiv:2605.06240v2 Announce Type: replace-cross
Abstract: Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inheri...
By Amirhossein Yousefiramandi
arXiv:2607. 02893v1 Announce Type: new Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width.
By Hamish Ogilvy
The study re‑examines a reported advantage of a routed ternary (1.58‑bit) language model over a full‑precision transformer at 60K parameters. By running controlled experiments with multiple seeds and a fixed training recipe, the authors find that the apparent benefit largely stems from the choice of baseline model shape rather than the ternary architecture itself. While the routed model does outperform other shapes at a larger 130M‑byte budget, its advantage diminishes when a plain gated diagonal‑SSM block is used, and the ternary penalty varies with architecture and quantization details.
By Gautam Veldanda
arXiv:2608. 13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable.
By Teng-Ruei Chen
arXiv:2608. 12026v1 Announce Type: new Abstract: Post-training quantization pipelines routinely leave the softmax output layer in high precision.
By Joao V. Cavalcanti, Ashia C. Wilson
arXiv:2608. 06564v2 Announce Type: replace Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts.
By Zekun Wu, Swati Dhiman, Adriano Koshiyama
The paper presents a 4‑bit quantization recipe, Minima: NVFP4 W4A4, that fully quantizes all linear layers—including the Gated DeltaNet (GDN) recurrent blocks—of the 27‑billion‑parameter Qwen3.8 LLM. Across a suite of benchmarks (perplexity, MMLU‑Pro, GSM8K, AIME'25, GPQA‑Diamond, LiveCodeBench, and RULER retrieval), the quantized model matches BF16 performance within seed noise while being 17.5 GiB in size and 14–19 % faster at prefill. The authors attribute this success to four mechanisms: block‑scaling of residuals, robust gate projections, the delta‑rule recurrence’s noise‑plateau behavior, and the per‑token quantization cost’s dilution over long contexts.
By Sergii Kozyrev, Davyd Maiboroda