arXiv Machine Learning By Eric Fock

Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample

Read the original on arXiv Machine Learning →

The paper introduces a reference‑free instrument that, from a single fit and without an oracle, can detect whether a hybrid PDE‑parameter estimator’s assumed operator is misspecified and distinguish this from mere parameter unidentifiability. In a self‑adjoint parabolic inverse problem, the proposed information‑matrix statistic correctly identifies misspecification with low false‑positive rates, while remaining silent when the design is correctly specified but non‑identifiable. The study demonstrates that conventional accuracy checks can miss significant operator errors, and it maps out the instrument’s blind spots and conditions under which its guarantees hold.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The study investigates how the composition of data during the mid‑training phase of language models affects performance across multiple domains. Experiments with Qwen3‑8B‑Base on five distinct KOR‑Bench domains show that moderate coverage (10%‑40%) yields the best per‑domain results, and that alignment passes cannot fully close the performance gaps created by mid‑training data choices. Additionally, zero coverage in mid‑training severely degrades accuracy, while a carefully tuned allocation can provide the largest overall pipeline improvement.

By Yunpeng Xu, Kun Zheng
arXiv Machine Learning
5d ago

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization

The paper argues that reporting scale‑invariant statistics such as cosine similarity without their noise floor is misleading when evaluating interpretability transfer from full‑precision to quantized neural networks. It derives a closed‑form expression for the expected cosine similarity based on a dimensionless parameter κ = n ho^2/d, measures the class separation ρ on real activations, and shows that a reported cosine of 0.996 between full‑precision and INT4 models cannot be interpreted as preservation without knowing the sample size n. The authors demonstrate that at INT4 the direction of the interpretability artifact rotates beyond the estimator’s own noise, while at INT8 no significant movement is detected, and they highlight that scale‑invariant metrics cannot distinguish between translation and attenuation of a transferred decision variable.

By Pranav Varshney