arXiv Machine Learning By Alina Kapanova, Arun Kanhai, Natan Vidra, Spurthi Setty

AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents

Read the original on arXiv Machine Learning →

arXiv:2608. 00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 3

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

arXiv:2607. 29577v1 Announce Type: new Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once.

By Ismayil Ismayilov, Atakan Kara, Kaan Oktay