Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Related stories
EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning
arXiv:2607. 01789v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale efficiently but remain costly to adapt due to redundant experts and uniform parameter allocation.
Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling
The paper investigates converting a large pretrained transformer (1.4âŻB parameters) into a smaller sibling (410âŻM) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a lowâbudget, structureâaware compensationâseparating leastâsquares function alignment from varianceâpreserving rescalingâyields significant gains on tokenâefficient training, outperforming subcloning and standard distillation pipelines at matched budgets.
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.
20x Faster TRL Fine-tuning with RapidFire AI
How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws
The paper proposes a theoretical framework for scheduling highâquality data in large language model training by extending functional scaling laws to account for timeâvarying data quality. It identifies two regimesânoiseâlimited and signalâlimitedâwhere highâquality data should be used differently, and introduces a DropâStableâRampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
arXiv:2606. 02559v1 Announce Type: cross Abstract: Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules.
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
The paper introduces LOOM, a method for scaling looped mixtureâofâexperts (MoE) Transformers beyond the typical twoâloop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and reâinjecting the input embedding, and it prevents expert selection collapse by using perâloop routers and a looping residual to maintain computational diversity. Experiments on 100âŻMâ1.7âŻB parameter models show stable scaling to 9â12 loops, with significant perplexity reductions and zeroâshot accuracy gains under nearâisoâFLOP conditions.
Sparse Layers are Critical to Scaling Looped Language Models
arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.
The Fine-Tuning Trap: Evaluating Negative Transfer and the Role of PEFT in Sub-1B Mathematical Reasoning
arXiv:2606. 06920v1 Announce Type: cross Abstract: Deploying Small Language Models (SLMs) on edge devices requires efficient fine-tuning strategies that adapt models to new tasks without degrading their general capabilities.
Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
arXiv:2607. 09172v1 Announce Type: cross Abstract: Large Language Models are reshaping how software is developed and maintained.
CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference
arXiv:2606. 10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass.