AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,300 stories · RSS feed

arXiv Machine Learning
Jul 2

Information-Regularized Attention for Visual-Centric Reasoning

arXiv:2607. 00434v1 Announce Type: cross Abstract: Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting after full-parameter instruction tuning.

By Guohao Sun, Xiaofang Wang, Yash Patel, Mengchen Liu, Zhiqiang Tao, Praveen Krishnan
arXiv AI
Jul 2

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

arXiv:2607. 00304v1 Announce Type: cross Abstract: The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N.

By Zewen Liu
arXiv Machine Learning
Jul 2

Explainability in mulimodal deep transformation models for stroke outcome prediction

arXiv:2504. 06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretability remains limited.

By Lisa Herzog, Jonas Br\"andli, Maurice Schneeberger, Loran Avci, Nordin Dari, Martin H\"ansel, Hakim Baazaoui, Pascal B\"uhler, Susanne Wegener, Beate Sick
arXiv Machine Learning
Jul 2

A Mechanistic View of Authority Hierarchy in LLM Sycophancy

arXiv:2607. 00415v1 Announce Type: cross Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence.

By Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen