The paper introduces a reference‑free instrument that, from a single fit and without an oracle, can detect whether a hybrid PDE‑parameter estimator’s assumed operator is misspecified and distinguish this from mere parameter unidentifiability. In a self‑adjoint parabolic inverse problem, the proposed information‑matrix statistic correctly identifies misspecification with low false‑positive rates, while remaining silent when the design is correctly specified but non‑identifiable. The study demonstrates that conventional accuracy checks can miss significant operator errors, and it maps out the instrument’s blind spots and conditions under which its guarantees hold.
By Eric Fock
arXiv:2606. 25151v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) embed governing equations in their loss function, enabling mesh-free solutions to partial differential equations.
By David McShannon, Nicholas Dietrich
arXiv:2607. 04113v1 Announce Type: new Abstract: Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $\sigma_{\min}$, at which the score is stiff and the flow develops a boundary layer.
By Shiheng Zhang
arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.
By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2609.07162v1 Announce Type: new
Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
By Xin Xu
arXiv:2609.26272v1 Announce Type: new
Abstract: Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to...
By Jian Xu
arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.
By Sunny Dubey
arXiv:2607. 11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck.
By Jeff Helzner
The paper investigates a failure mode in Graph-JEPA, a joint‑embedding predictive model trained on a large scientific‑reasoning graph. Despite achieving high linear‑probe accuracy and effective rank, the learned representation contains almost no usable instance information, as shown by retrieval metrics. The authors diagnose the issue to variance allocation in the objective, propose a repair that restores near‑perfect information recovery, and demonstrate that the problem persists even after repair, highlighting limitations in the evaluation metrics used.
By Gollam Rabby, S\"oren Auer