Goal reasoning in Non-Axiomatic Logic (NAL) explains how an adaptive system derives means for realizing desired events under insufficient knowledge and resources. However, the representation of avoidance is less clear.
arXiv:2606. 31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones.
By Taeyoun Kim, Aviral Kumar
The article discusses goals as cognitive states that combine with world knowledge to guide purposeful behavior, emphasizing their compositional nature and relation to rational action. It draws parallels between goal representations and the syntax‑semantics interface in linguistics and logic, highlighting questions about expressivity, design, and efficiency of different goal languages. The authors synthesize research on goal representation properties, propose a broader design space, and suggest that distinguishing form and meaning can clarify assumptions, inform cognition‑motivation interactions, and identify variation axes in goal conceptions.
By David M. Abel, Mark K. Ho
arXiv:2608. 15673v1 Announce Type: cross Abstract: Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those facts imply under a given policy.
By Satchit Chatterji, Shihan Wang, Giovanni Sileno, Erman Acar
arXiv:2608.30197v1 Announce Type: new
Abstract: Safety alignment is essential for deploying large language models, requiring systems to prevent harmful compliance while preserving helpfulness on beni...
By Hoejoon Kwon, Byeonggeuk Lim, Kahyeon Kim, YoungBin Kim
arXiv:2510. 15395v2 Announce Type: replace Abstract: An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates.
By Rubi Hudson