AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,089 stories · RSS feed

arXiv Machine Learning
Jul 7

Do ECG Foundation Models Transfer to Rare Cardiac Diseases? Evidence from Brugada Syndrome Detection

arXiv:2607. 03009v1 Announce Type: new Abstract: Background: Foundation models (FMs) trained on large-scale unlabeled physiological data have emerged as a promising paradigm for medical artificial intelligence.

By Beatrice Zanchi, Giuliana Monachino, Alvise Dei Rossi, Luigi Fiorillo, Georgia Sarquella-Brugada, Giulio Conte, Francesca Dalia Faraci
arXiv Machine Learning
Jul 7

Adversarial LassoNet: Robust Feature Selection via Stability-Driven Sparse Learning

arXiv:2607. 03839v1 Announce Type: new Abstract: Sparse feature selection is critical for high-dimensional machine learning, yet traditional $\ell_1$-regularized methods are often brittle under observational noise and spurious correlations, leading to unstable feature supports and degraded generalization.

By Zhen Huang, Peicheng Xu, Junbiao Pang, Yulong Zheng
arXiv Machine Learning
Jul 7

CSympNet-ID: conformal-symplectic map learning for linearly damped Hamiltonian systems

arXiv:2607. 03339v1 Announce Type: new Abstract: Learning dissipative dynamics from discrete observations is essential for reliable long-horizon prediction and physically meaningful parameter identification.

By Jiale Gong (School of Mathematics), Pengzhan Jin (National Engineering Laboratory for Big Data Analysis and Applications, Peking University, Beijing, China), Dongyang Kuang (School of Mathematics), Lu Li (School of Mathematics), Yifa Tang (State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China)
arXiv AI
Jul 7

Multi-Way Representation Alignment

arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.

By Akshit Achara, Tatiana Gaintseva, Mateo Mahaut, Pritish Chakraborty, Viktor Stenby Johansson, Melih Barsbey, Emanuele Rodol\`a, Donato Crisostomi
arXiv Machine Learning
Jul 7

Seeing Through WiFi: Lightweight Human Pose Estimation with Dynamic Kernel Attention

arXiv:2607. 03196v1 Announce Type: cross Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devices while preserving individual privacy concerns.

By Toan D. Gian, Van-Dinh Nguyen, Vo Phi Son, Nhan Thanh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Nguyen Cong Luong, Symeon Chatzinotas