AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,629 stories · RSS feed

arXiv AI
Aug 11

Scaling Inherently Interpretable Language Models

arXiv:2608. 07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish.

By Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo
arXiv AI
Aug 11

Human-Guided Causal Knowledge Injection for Virtual Cells

arXiv:2608. 08430v1 Announce Type: cross Abstract: Virtual cells employ machine learning models to simulate and predict cellular behaviors, serving as a critical computational framework for investigating health and disease.

By Pengcheng Wang, Changjian Chen, Zhuo Tang, You Wu, Long Wang, Feng Yu, Kenli Li