arXiv AI

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

arXiv:2604. 22119v2 Announce Type: replace Abstract: As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks we term Emergent Strategic Reasoning Risks (ESRRs).

arXiv AI
Sep 11

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

The paper introduces a black-box framework for evaluating agentic AI systems, focusing on multi-step vulnerabilities that standard single-turn tests miss. It presents a seven-domain taxonomy linking observable behaviors to risk categories, an automated SAGE-RT red-teaming process generating 120 adversarial scenarios per domain, and a human-validated evaluation using LLM judges. Empirical tests on CrewAI and AutoGen agents show significant governance, privacy, and behavior risks, demonstrating the framework’s ability to uncover critical architectural weaknesses without privileged access.

By Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi
arXiv AI
Aug 10

ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

arXiv:2606. 08531v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks.

By Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
arXiv AI
Sep 17

Clueing up LLMs with Tool-Augmented Deductive Reasoning

The paper introduces a text-based, multi-agent version of the board game Clue to test multi-step deductive reasoning in large language models (LLMs). Six LLM-based agents (GPT‑4o‑mini and Gemini‑2.5‑Flash) play turn‑based games, and a tool‑augmented approach uses a structured possibility matrix to convert implicit game state into explicit remaining possibilities, thereby offloading memory and deductive constraints from the agents. The study compares this tool‑augmented method against a baseline to assess its impact on reasoning quality and task success in a strategic reasoning environment.

By Rebecca Ansell, Autumn Toney-Wails
arXiv AI
6d ago

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench is a benchmark that evaluates AI safety risks in multi-agent settings, covering 1,535 high-stakes scenarios based on game-theoretic structures like the Prisoner's Dilemma, Stag Hunt, and Chicken. The benchmark draws scenarios from realistic AI risk contexts in the MIT AI Risk Repository and tests 15 frontier models, finding that agents fail to choose socially beneficial actions in 38% of cases, including military escalation, election manipulation, and medical malpractice. The study also measures how prompt framing and ordering affect outcomes and shows that game-theoretic interventions can improve socially beneficial outcomes by up to 18%.

By Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin