arXiv Machine Learning
Aug 19

Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample

The paper introduces a reference‑free instrument that, from a single fit and without an oracle, can detect whether a hybrid PDE‑parameter estimator’s assumed operator is misspecified and distinguish this from mere parameter unidentifiability. In a self‑adjoint parabolic inverse problem, the proposed information‑matrix statistic correctly identifies misspecification with low false‑positive rates, while remaining silent when the design is correctly specified but non‑identifiable. The study demonstrates that conventional accuracy checks can miss significant operator errors, and it maps out the instrument’s blind spots and conditions under which its guarantees hold.

By Eric Fock
arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard