arXiv:2608. 05710v1 Announce Type: new Abstract: When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes.
By Keziah Naggita
The article discusses autonomous systems as the pinnacle of AI development, emphasizing the need to blend connectionist and symbolic AI within systems engineering. It introduces a generic agent architecture that organizes behavior around long‑term memory and outlines challenges in linking sensory data to structured memory, goal‑oriented decision making, planning, and agent coordination for collective intelligence. The authors also explore agent trustworthiness, noting it extends beyond behavior to include cognitive validity, and propose methods for its evaluation while highlighting the gap between current capabilities and the envisioned autonomous multi‑agent systems.
By Joseph Sifakis
arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
By Ranjan Mishra, Jakob Schoeffer
One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve developed an algorithm which can infer what humans want by being told which of two proposed behaviors is better.
arXiv:2601. 06077v2 Announce Type: replace-cross Abstract: This work aims to rigorously define the values of perception, prediction, communication, and common sense in decision making.
By Aolin Xu
We (along with researchers from Berkeley and Stanford) are co-authors on today’s paper led by Google Brain researchers, Concrete Problems in AI Safety. The paper explores many research problems around ensuring that modern machine learning systems operate as intended.