arXiv Machine Learning By Kamil Ciosek, Nicol\`o Felicioni, Juan Elenter, Ehsan Imani

Gradient Prediction with Control Variates in the Cheap-Forward Regime

Read the original on arXiv Machine Learning →

The paper investigates whether idle inference resources can help cut the high cost of scarce GPU usage during training. Using a simulated compute ledger that bills fleet work at a fraction of a GPU forward pass, the authors propose an algorithm that predicts gradients with a low‑precision, inference‑style reverse‑mode program and then refines these predictions with a few exact gradients via a control variate, turning approximation error into variance rather than bias. Experiments on a 124‑million‑parameter language model and across models ranging from 10 M to 774 M parameters show that the method can reduce simulated ledger cost when fleet work is cheap, though it also exhibits both successful transfers and failures, and does not evaluate inference‑only hardware or full optimizer‑by‑batch‑size sweeps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 16

OPEN-1B: A Fully Auditable Training Run

The paper introduces Open-1B, a language model trained under a new fully auditable regime that ensures every training operation is reproducible on heterogeneous commodity hardware with bitwise certainty. By enforcing a fixed order on sources of nondeterminism—GPU reductions, data batch ordering, and inter/intra-node communication—the authors enable auditors to replay and verify individual training steps on a single machine. The release includes the full pretraining dataset, all intermediate checkpoints, the training codebase, and an audit harness for step-by-step verification.

By John Donaghy, Brian Wilcox, O\u{g}uzhan Ersoy, Shikhar Rastogi, Adam St Arnaud, Alexey Titov, Jordan Greenberg, Ben Fielding, Harry Grieve