AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,629 stories · RSS feed

arXiv Machine Learning
Aug 7

Phylogenetic Tree Inference with Tropical Axial Attention

arXiv:2605. 13894v2 Announce Type: replace-cross Abstract: In this work, we introduce a Tropical Axial Attention neural reasoning architecture that replaces vanilla softmax dot-product attention with max-plus operators, inducing a piecewise-linear structure aligned with dynamic programming formulations.

By Chris Teska, Kurt Pasque, Ruriko Yoshida, Baran Hashemi
arXiv AI
Aug 7

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

arXiv:2608. 06020v1 Announce Type: new Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes.

By Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong