arXiv AI By Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto

Rules or Character? Scaling Laws for AI Safety Design

Read the original on arXiv AI →

arXiv:2608. 13345v1 Announce Type: new Abstract: Artificial Intelligence (AI) safety systems combine character shaping (e.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 20

Harmonizing AI Safety Thresholds

arXiv:2607. 16112v1 Announce Type: new Abstract: Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies.

By Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey
arXiv AI
Jul 7

Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control

arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.

By Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum