The paper proposes a theory for judging post-hoc debates in AI, focusing on properties like reproducibility, robustness, groundedness, and explainability. It evaluates two debate‑judgement methods—LLM judges and formal computational argumentation semantics—finding similar accuracy but noting that argumentation semantics offers stronger formal guarantees. The study suggests that argumentation semantics is a preferable framework for principled debate judges in AI systems.
By Xiang Yin, Adam Dejl, Antonio Rago, Lihu Chen, Francesca Toni
We (along with researchers from Berkeley and Stanford) are co-authors on today’s paper led by Google Brain researchers, Concrete Problems in AI Safety. The paper explores many research problems around ensuring that modern machine learning systems operate as intended.
We’ve written a paper arguing that long-term AI safety research needs social scientists to ensure AI alignment algorithms succeed when actual humans are involved. Properly aligning advanced AI systems with human values requires resolving many uncertainties related to the psychology of human rationality, emotion, and biases.
arXiv:2606. 03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identification.
By Sanjay Das, Ran Elgedawy, Ethan Seefried, Ryan Burchfield, Tirthankar Ghosal
arXiv:2506. 13609v2 Announce Type: replace Abstract: Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks.
By Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras, Lijie Chen, Jiawei Li, Zhiyang Xun
Ensuring that AI systems are built, deployed, and used safely is critical to our mission.
OpenAI introduces CoT-Control and finds reasoning models struggle to control their chains of thought, reinforcing monitorability as an AI safety safeguard.
The article "Math for AI safety: an invitation for mathematicians" calls for new mathematical tools to ensure AI remains understandable, controllable, and cooperative. It outlines specific mathematical fields—logic, game theory, probability, algebra, representation theory, analysis, and geometry—each paired with an open problem tailored for mathematicians without AI safety background. The piece invites researchers to contribute to designing AI that is legible, steerable, and aligned with human values.
By Lionel Levine
The paper introduces Safety Nudges, a browser-based tool that displays lightweight, in situ flags when a conversational AI exhibits risky behavior such as hallucination or overconfidence. In a two‑week field study with 45 frequent chatbot users, participants reported that the nudges were useful, clear, and minimally disruptive, and most felt more aware of potential AI harms. However, increased awareness did not automatically translate into measurable changes in user behavior, underscoring the need for relevance, calibration, and user control in nudge design.
By Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
OmouAI is an interactive deliberation system that combines large language models with computational argumentation to facilitate policy debates involving humans and simulated personas such as stakeholders, experts, or devil’s advocates. Each persona generates its own arguments, which are assembled into a shared argumentation framework that users can contest, add to, or revise, ensuring human oversight. The system evaluates arguments using deterministic argumentative semantics against external goals like the UN Sustainable Development Goals, providing faithful explanations and indicating how policy recommendations affect those goals.
By Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto
arXiv:2510. 26518v2 Announce Type: replace Abstract: Human feedback is critical for aligning AI systems to human values.
By Rishub Jain, Sophie Bridgers, Lili Janzer, Rory Greig, Tian Huey Teh, Vladimir Mikulik