A JEPA Recipe for Tabular Foundation Models
arXiv:2609.25541v1 Announce Type: new Abstract: Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (Le...
The paper introduces a stability theory for Joint-Embedding Predictive Architectures (JEPAs), showing that training dynamics involve a driving force and a decay effect that determine representation collapse. By linearising the gradient flow, the authors derive a per‑mode stability ratio that separates data‑side and predictor‑side contributions, predicting a phase boundary confirmed across 800 configurations. Using this insight, they propose ResidualPred, a transformer predictor that biases attention toward the identity at initialization, improving representation rank and downstream accuracy on tabular and image benchmarks.
arXiv:2609.25541v1 Announce Type: new Abstract: Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (Le...
The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.
arXiv:2604. 00230v2 Announce Type: replace Abstract: Neural collapse (NC) -- the convergence of penultimate-layer features to a simplex equiangular tight frame -- is well understood at equilibrium, but the dynamics governing its onset remain poorly characterised.
arXiv:2605. 10840v3 Announce Type: replace-cross Abstract: We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories.
arXiv:2606. 00091v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) have reshaped self-supervised representation learning in vision.
arXiv:2607. 15449v1 Announce Type: new Abstract: Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator.
arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
arXiv:2608. 15483v1 Announce Type: new Abstract: Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers.
The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.
Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-h...
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.