AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,089 stories · RSS feed

arXiv Machine Learning
Jul 3

Fast and Accurate Anomaly Detection in Time Series

arXiv:2607. 02046v1 Announce Type: new Abstract: Anomaly detection is a critical and evolving field in Machine Learning, with applications targeting different domains such as cybersecurity, finance, healthcare, manufacturing and IoT (Internet of Things) systems.

By Emanuele Mele, Massimo Cafaro, Angelo Coluccia, Italo Epicoco
arXiv AI
Jul 3

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

arXiv:2511. 05150v2 Announce Type: replace-cross Abstract: Molecular biomarker testing in pathology is often costly and tissue-consuming, limiting scalable clinical deployment.

By Jingsong Liu, Han Li, Zhengyang Xu, Franz-Leonard Klaus, Fabian St\"ogbauer, Shihui Zu, Weiwei Zhou, Atsuko Kasajima, Felix Schicktanz, Alexander Muckenhuber, Julius Shakhtour, Jiale Yu, Tiannan Zheng, Xun Ma, Maggie Wang, Christian Grashei, Bao Li, Guiyang Jiang, Hongming Xu, Shaohua Kevin Zhou, Nassir Navab, Peter J. Sch\"uffler
arXiv AI
Jul 3

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

arXiv:2607. 01934v1 Announce Type: cross Abstract: This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K-12.

By Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana, Alvaro Ortigosa, Enrique Blas, Aythami Morales
arXiv AI
Jul 3

Generative AI and Federated Learning for Intrusion Detection Systems: A Survey

arXiv:2607. 01305v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments.

By Jiefei Liu, Abu Saleh Md Tayeen, Pratyay Kumar, Qixu Gong, Wenbin Jiang, Huiping Cao, Satyajayant Misra, Jayashree Harikumar
arXiv AI
Jul 3

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

arXiv:2607. 01336v1 Announce Type: cross Abstract: Neural Quantum States (NQS) are a remarkably expressive class of variational ans\"atze for quantum many-body wavefunctions, yet little is understood about their internal mechanisms: trained on variational objectives alone, how do NQS accurately capture physical observables that they have never been explicitly optimized for?

By Zihao Qi, Christopher Earls