arXiv Machine Learning

How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?

The study investigates how distribution shift influences the benefits of pretraining neural PDE surrogates for computational fluid dynamics. Researchers pretrained a model on 254,909 RANS solutions from one airfoil family and fine‑tuned it on a new family under two target settings—identical Spalart‑Allmaras (SA) modeling and SA with added $e^N$ transition modeling—while keeping freestream ranges matched. Results show that at 1,000 fine‑tuning samples, the pretrained model matches a from‑scratch model trained on 3.25× more data for the same‑SA target and 2.58× more for the transition‑modeled target; by 5,000 samples the advantage reverses. Additionally, increasing the number of distinct airfoils in the fine‑tuning set reduces error for both targets, but the improvement is significant only for the same‑SA case.

arXiv AI
Sep 25

Revalidation Beats Stateful Routing for Scientific Surrogates Under Distribution Shift

The study introduces RegimeShift‑Surrogates, a streaming benchmark that tests surrogate models across eight tasks and multiple regimes. It compares revalidation—choosing the model with lowest current‑window validation loss—to stateful adaptive controllers and finds that revalidation consistently outperforms stateful methods, achieving lower mean log regret in most task‑scenario combinations. The results suggest that fresh validation evidence is more valuable than carrying over past evidence when dealing with distribution shifts.

By Harshil Lodhiya
arXiv Machine Learning
Jul 9

Optimization-Embedded Active Multi-Fidelity Surrogate Learning for Multi-Condition Airfoil Shape Optimization

arXiv:2603. 17057v2 Announce Type: replace-cross Abstract: Active multi-fidelity surrogate modeling is developed for multi-condition airfoil shape optimization to reduce high-fidelity CFD cost while retaining RANS-consistent aerodynamic metrics.

By Isaac Robledo, Alberto Vilari\~no, Arnau Mir\'o, Oriol Lehmkuhl, Carlos Sanmiguel Vila, Rodrigo Castellanos
arXiv Machine Learning
Sep 7

A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations

The paper introduces a correction framework that grounds a CFD-trained deep learning surrogate model for aerospace aerodynamics using wind‑tunnel pressure‑sensor (PSP) data. By training a correction network on spatially registered PSP measurements at two Mach numbers, the authors adjust the surrogate’s predictions without retraining its core parameters, achieving improved agreement with experimental pressure distributions—especially at the wing suction peak and shock location. The grounded surrogate matches measurements within 2.3–2.7% of the Cp range on unseen angles of attack and outperforms simple interpolation between measured states.

By Nitin Nagesh Kulkarni, Dheeraj Vemula, Yin Yu, Peter Lyu, Juan J. Alonso
arXiv Machine Learning
1d ago

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.

By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran
arXiv AI
Sep 16

Schema-Adaptive Action-Conditioned JEPA for Cross-Machine CNC Transfer under Partial Sensor Overlap

The paper introduces a schema‑adaptive action‑conditioned Joint‑Embedding Predictive Architecture (SAAC‑JEPA) for cross‑machine CNC transfer when only a subset of sensors overlap between source and target machines. Experiments show that pretraining does not improve source‑only forecasting, but a carefully selected action‑conditioned JEPA model achieves a zero‑shot RMSE of 0.546 on the target, outperforming persistence but falling short of certain baseline models. Ablation studies reveal that adding RevIN improves RMSE but harms calibration, and limited post‑lock adaptation can further reduce error.

By Ayoub Louaye Bouaziz, Matthieu Ostertag, Anton Demasles
arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard
arXiv Machine Learning
Aug 28

Curating Same-Family Neural Networks for LLM-Guided Model Improvement: A Controlled Case Study

The study investigates whether a curated same-family neural network experiment can guide large language model (LLM)-based improvements for a low-performing target model under equal generation and evaluation budgets. Using TuneNNGen, an extension of NNGPT, the authors compare source-guided generation with target-only generation on CIFAR-10, SVHN, Imagenette, and CIFAR-100 datasets, achieving significant accuracy gains across these benchmarks. The results demonstrate that the benefits depend on source-target compatibility and LLM adaptation, rather than merely on stored source accuracy.

By Kabir Dev Paul Baghel, Radu Timofte, Dmitry Ignatov
arXiv AI
Jun 26

Error-Conditioned Neural Solvers

arXiv:2606. 27354v1 Announce Type: cross Abstract: Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution.

By Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak, Seungryong Kim, Brian Bell, Jeong Joon Park