AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,195 stories · RSS feed

arXiv Machine Learning
Jul 28

Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization

arXiv:2602. 03570v2 Announce Type: replace Abstract: Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space.

By Bixing Wu, Yuhong Zhao, Zongli Ye, Jiachen Lian, Xiangyu Yue, Gopala Anumanchipalli
arXiv AI
Jul 28

Invariant Discovery for Networked Systems

arXiv:2607. 22944v1 Announce Type: cross Abstract: Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing them by hand demands rare expertise in both formal logic and networking.

By Hongyu H\`e, Alexander Krentsel, Sylvia Ratnasamy, Maria Apostolaki
arXiv Machine Learning
Jul 28

Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

arXiv:2605. 07407v2 Announce Type: replace Abstract: We show that information can be transferred post-hoc across independently trained health foundation models (FMs), each pretrained on ~20M minutes of wearable sensor data from ~172K participants, by aligning their data-dependent coordinate systems.

By Gajendra Katuwal, Advait Koparkar, Salar Abbaspourazad, Anshuman Mishra, Sarvesh Kirthivasan
arXiv AI
Jul 28

EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

arXiv:2607. 24553v1 Announce Type: cross Abstract: Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings.

By Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong
Hugging Face Trending Papers
Jul 27

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change.