arXiv AI By Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe, Hal Daum\'e III

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

Read the original on arXiv AI →

arXiv:2606. 07631v1 Announce Type: cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

Leakage-Free Evaluation and Distribution-Robust Spatio-Temporal Graph Learning for Inductive Kriging

The paper introduces a leakage‑free 3×3 spatio‑temporal partition for evaluating inductive kriging, ensuring training, validation, and testing occur on distinct spatial and temporal domains. It proposes DRIK, a framework that includes Spatial Continuity Regularization, Masked Flow Disambiguation, and Structural Domain Expansion to mitigate structural shifts from unseen nodes. Experiments on six datasets show DRIK outperforms existing baselines, reducing MAE by up to 12.48% and achieving lower test‑to‑validation MAE ratios under the stricter evaluation protocol.

By Chen Yang, Changhao Zhao, Haoyang Zhao, Youquan He, Chen Wang, Jiansheng Fan
arXiv Machine Learning
Sep 2

Auditing Frozen-Encoder Anomaly Detection Across Mechanical Systems: Representation Provenance, Calibration, and Protocol Effects

This paper presents a reproducibility audit of frozen‑encoder anomaly detection experiments originally reported on arXiv. The authors confirm that the numerical discrimination results can be reproduced from the preserved artifacts, but they find that the claimed causal link to interferometric pretraining is unsupported. They show that near‑zero embeddings and architectural choices, rather than a morphological prior from gravitational‑wave instrumentation, explain the observed anomaly‑detection performance.

By Jose S\'anchez Andreu