AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,037 stories · RSS feed

arXiv Machine Learning
Jul 31

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv:2607. 27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses.

By Yixuan Duan, Wei Qiu
arXiv Machine Learning
Jul 31

ZAPs: A Reward Attribution Framework for DeFi Ecosystems with Adversarial-Robust Scoring via Parallel Anomaly Ensemble Detection

arXiv:2607. 27859v1 Announce Type: cross Abstract: Incentive programs are central to user acquisition in decentralized finance, but many reward systems rely on raw volume, transaction count, and wallet count, making them vulnerable to bots and sybil operations.

By Girish G N, Ashutosh Sahoo, Ajay Bhat, Akshay SP, Gurukiran S, Parag Paul, Dhanashekar Kandaswamy
arXiv Machine Learning
Jul 31

Robust Wavelength Selection for Partial Least Squares Sugar Content Estimation Using Combinatorial Bayesian Optimization

arXiv:2607. 27645v1 Announce Type: cross Abstract: Wavelength selection is one of the important preprocessing methods in near-infrared spectroscopy to improve prediction accuracy and interpretability of spectral data.

By Mitsunobu Kanebako, Ami S. Koshikawa, Masaru Hitomi, Takuro Tanaka, Mahito Chiba, Maiko Mori, Masayuki Ohzeki
arXiv AI
Jul 31

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

arXiv:2607. 26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction.

By Xinyu Yang, Tianxing Chen, Honghao Su, Minxuan Wang, Chenze Yu, Zhangzheng Tu, Yue Chen, Yuxiao Huo, Lingfeng Zhang, Yan Huang, Yan Qin, Shaolong Zhu, Qiwei Liang, Hekun Tian, Shujia Liu, Guangyu Chen, Junhao Gong, Zixuan Li, Wenwei Lin, Zijian Lin, Wenxuan Zhu, Eric J Chen, Yue Yuan, Qize Yu, Jiaqi Liang, Haowen Yan, Hengfei Zhao, Weijie Wan, Zikun Xiao, Junyuan Tang, Baijun Chen, Kai-Chong Lei, Kaixuan Wang, Kailun Su, Zanxin Chen, Yao Mu, Renjing Xu, Chuqiao Lyu, Qi Xiong, Ping Luo, Wenbo Ding
arXiv AI
Jul 31

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

arXiv:2607. 26060v1 Announce Type: cross Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment.

By Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi