arXiv AI

Schema-Adaptive Action-Conditioned JEPA for Cross-Machine CNC Transfer under Partial Sensor Overlap

The paper introduces a schema‑adaptive action‑conditioned Joint‑Embedding Predictive Architecture (SAAC‑JEPA) for cross‑machine CNC transfer when only a subset of sensors overlap between source and target machines. Experiments show that pretraining does not improve source‑only forecasting, but a carefully selected action‑conditioned JEPA model achieves a zero‑shot RMSE of 0.546 on the target, outperforming persistence but falling short of certain baseline models. Ablation studies reveal that adding RevIN improves RMSE but harms calibration, and limited post‑lock adaptation can further reduce error.

arXiv AI
Sep 24

World Models for Cross-Machine CNC Transfer under Partial Sensor Overlap

The paper investigates whether a command‑conditioned latent world model, trained on a source CNC machine with 17 sensor channels, can transfer its predictive capability to a target machine that shares only 10 of those channels. Experiments show that latent‑predictive pretraining offers no advantage over training from scratch on the source data, and that the transferred model outperforms a persistence baseline on the target but falls short of forecasters that normalize each input window by its own statistics. The study highlights that cross‑machine transfer under partial sensor overlap presents a unique challenge for command‑conditioned world models.

By Ayoub Louaye Bouaziz, Matthieu Ostertag, Anton Demasles
arXiv AI
Jun 18

Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories

arXiv:2605. 10840v3 Announce Type: replace-cross Abstract: We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories.

By Yixuan Yang, Mehak Arora, Ryan Zhang, Baraa Abed, Junseob Kim, Tilendra Choudhary, Md Hassanuzzaman, Kevin Zhu, Ayman Ali, Chengkun Yang, Alasdair Edward Gent, Victor Moas, Rishikesan Kamaleswaran
arXiv Machine Learning
Sep 2

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

The paper investigates when online adaptation benefits edge time‑series forecasting under distribution drift, using a leakage‑free streaming protocol on six public multivariate datasets. It shows that the warmup budget for static baselines and the choice of learning rate can bias perceived adaptation gains, and that a validation‑only procedure selecting warmup and optimizer rates yields Adam outperforming SGD with momentum in most settings. The study also examines accuracy versus adaptation‑state memory and per‑update latency for different adaptation strategies, highlighting parameter‑efficient variants that are nondominated on the memory axis.

By Takumi Fujimoto, Hiroaki Nishi
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard
arXiv Machine Learning
Jul 2

LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning

arXiv:2607. 00958v1 Announce Type: new Abstract: Time series are central to modern data mining applications, from industrial telemetry and server metrics to finance and physiology, yet time-series self-supervised learning often depends on view and augmentation choices that encode domain-specific invariances.

By Alexander Chemeris, Ming Jin, Randall Balestriero
arXiv AI
6d ago

Revalidation Beats Stateful Routing for Scientific Surrogates Under Distribution Shift

The study introduces RegimeShift‑Surrogates, a streaming benchmark that tests surrogate models across eight tasks and multiple regimes. It compares revalidation—choosing the model with lowest current‑window validation loss—to stateful adaptive controllers and finds that revalidation consistently outperforms stateful methods, achieving lower mean log regret in most task‑scenario combinations. The results suggest that fresh validation evidence is more valuable than carrying over past evidence when dealing with distribution shifts.

By Harshil Lodhiya