arXiv Machine Learning By O. M. Kiselev

Boundary-Aware Quantization: Finite-Scale Decision Geometry of Neural Classifiers

Read the original on arXiv Machine Learning →

arXiv:2607. 01478v1 Announce Type: cross Abstract: We measured quantization-induced decision-boundary changes using local logit-margin radii, first-order boundary displacement, normal variation, slice-boundary Jaccard distance, grid prediction changes, multiclass junction counts, and low-margin boundary-band flips.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

By Yunpeng Xu, Kun Zheng
arXiv Machine Learning
Sep 22

Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization

The paper investigates whether sink-aware attention head selection remains valid after 4‑bit NF4 weight‑only post‑training quantization. Using Sink Topology Consistency metrics, it finds that global rank preservation stays high across Qwen2.5 and Llama‑3.2 models, yet top‑k head overlap drops to 61–79% and layer‑specific sink‑mass shifts can be substantial. The study also shows that cross‑domain calibration degrades more than within‑domain precision and that recalibration with a small number of samples can recover most of the stability, though full‑map stability may require updating more layers.

By Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou
Hugging Face Trending Papers
Sep 8

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The paper investigates how per-domain data composition during the mid‑training phase (between pre‑training and alignment) affects model performance. Experiments with Qwen3‑8B‑Base across five KOR‑Bench domains show that a moderate coverage band (10%‑40%) yields the best performance for each domain, and that alignment passes cannot fully close the gaps created by suboptimal mid‑training allocations. Additionally, zero coverage during mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard