arXiv AI By Yunpeng Xu, Kun Zheng

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

Read the original on arXiv AI →

The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 8

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The paper investigates how per-domain data composition during the mid‑training phase (between pre‑training and alignment) affects model performance. Experiments with Qwen3‑8B‑Base across five KOR‑Bench domains show that a moderate coverage band (10%‑40%) yields the best performance for each domain, and that alignment passes cannot fully close the gaps created by suboptimal mid‑training allocations. Additionally, zero coverage during mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

arXiv AI
Aug 25

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.

By Pedro Santos
arXiv AI
6d ago

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula