Epistemic Norms for AI Safety and Alignment Research
arXiv:2607. 24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high.
arXiv:2606. 07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation.
arXiv:2607. 24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high.
arXiv:2608. 05656v1 Announce Type: cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies.
arXiv:2505. 22829v2 Announce Type: replace-cross Abstract: This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies.
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.
arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous,...
arXiv:2607. 07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires.
Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.
arXiv:2608. 00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction.
arXiv:2608.22160v1 Announce Type: new Abstract: Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at p...
arXiv:2607. 19292v1 Announce Type: cross Abstract: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios.
arXiv:2605. 02050v2 Announce Type: replace-cross Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies).
We show that ordinary business language --- "maximize profitability" --- induces profit-oriented ambiguity resolution: LLMs systematically dismiss ambiguous signals of potential safety violations to s...