arXiv:2608. 05710v1 Announce Type: new Abstract: When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes.
By Keziah Naggita
The article discusses autonomous systems as the pinnacle of AI development, emphasizing the need to blend connectionist and symbolic AI within systems engineering. It introduces a generic agent architecture that organizes behavior around long‑term memory and outlines challenges in linking sensory data to structured memory, goal‑oriented decision making, planning, and agent coordination for collective intelligence. The authors also explore agent trustworthiness, noting it extends beyond behavior to include cognitive validity, and propose methods for its evaluation while highlighting the gap between current capabilities and the envisioned autonomous multi‑agent systems.
By Joseph Sifakis
arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
By Ranjan Mishra, Jakob Schoeffer
One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve developed an algorithm which can infer what humans want by being told which of two proposed behaviors is better.
arXiv:2601. 06077v2 Announce Type: replace-cross Abstract: This work aims to rigorously define the values of perception, prediction, communication, and common sense in decision making.
By Aolin Xu
We (along with researchers from Berkeley and Stanford) are co-authors on today’s paper led by Google Brain researchers, Concrete Problems in AI Safety. The paper explores many research problems around ensuring that modern machine learning systems operate as intended.
The article discusses how artificial intelligence is beginning to automate scientific discovery, specifically in the realm of cognitive science. It outlines four key challenges for developing an automated science of the mind: representing experiments, generating synthetic behavior, synthesizing models, and closing the loop to discover psychological theories. The authors argue that addressing these challenges will enable AI to systematically advance our understanding of the mind.
By Akshay K. Jagadish, Milena Rmus, Kristin Witte, Marvin Mathony, Marcel Binz, Eric Schulz
arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.
By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
We’re proposing an AI safety technique called iterated amplification that lets us specify complicated behaviors and goals that are beyond human scale, by demonstrating how to decompose a task into simpler sub-tasks, rather than by providing labeled data or a reward function. Although this idea is in its very early stages and we have only completed experiments on simple toy algorithmic domains, we’ve decided to present it in its preliminary state because we think it could prove to be a scalable approach to AI safety.
arXiv:2606. 06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI?
By Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah, Catherine Arnett, Fazl Barez, Naomi Saphra
arXiv:2607. 06656v1 Announce Type: new Abstract: Machine learning models are often intended to augment rather than replace human decision makers, by providing information that is complementary to human judgement.
By Yewon Byun, Bryan Wilder
arXiv:2601.12053v2 Announce Type: replace-cross
Abstract: While foundation models have achieved remarkable results across a diversity of domains, they still rely on human-generated data, such as text...
By Ma\"el Donoso