arXiv Machine Learning By Yoshiyuki Ootani

Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization

Read the original on arXiv Machine Learning →

The paper introduces the fully‑enumerable transformer—a tiny transformer trained on tasks where every input can be evaluated exactly—as a scientific instrument for studying delayed generalization. It claims four unique capabilities: exact, falsifiable generalization ceilings; precise task surgery; direct observation of all weights; and survival‑time statistics that treat non‑grokking as censored data. A preregistered conservation study shows that two task‑side laws (a recoverability‑ceiling law and a role‑conflict delay law) hold across a 4,000‑fold increase in model size, while a weight‑decay law deforms predictably, demonstrating that the instrument can reveal lawful patterns of generalization across scales.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

The paper presents a causal analysis of a compressed VLA policy that performs well in offline tests but fails in closed‑loop execution on a simulated pick‑and‑place task. An 8‑layer distillation of Octo‑Base retains most parameters and passes all offline metrics, yet collapses during deployment, with early stages degrading gradually and final transport failing entirely. The failure is traced to a negative, late‑heavy residual in the action trace, and standard remedies (continued training, offline data, command‑level compensation, clamping) do not restore performance; only a minimal‑pair intervention that mixes deployment‑distribution rollouts with teacher data restores parity with the teacher. whyItMatters":"The study demonstrates that offline validation metrics alone are insufficient to guarantee closed‑loop success for compressed policies, highlighting the need for targeted deployment‑time testing and interventions."

By Fengze Jia (The Ohio State University)