arXiv AI

Learning Local Constraints for Reinforcement-Learned Content Generators

arXiv AI
3d ago

Conditional Generation of Creative Chess Puzzles with Diffusion Models

The paper presents a method for generating creative chess puzzles using masked diffusion models that can be conditioned on tactical themes and partial board positions. It introduces an auxiliary best‑move prediction task that boosts solution uniqueness by 11.6% and theme‑conditioning accuracy by 2.5%. A reinforcement learning framework further increases the yield of unique, theme‑matching puzzles by 89.1%, and the authors release open‑weights models for the community.

By Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi
arXiv AI
Jul 21

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization

arXiv:2607. 16206v1 Announce Type: new Abstract: This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the "Invisible Shackles" of mode collapse in Large Language Model (LLM) fine-tuning.

By Yujie Shen, Haowen Chen
arXiv AI
Jul 3

Evolutionary Wave Function Collapse

arXiv:2607. 02082v1 Announce Type: cross Abstract: Wave Function Collapse (WFC) is a widely used procedural content generation method that learns local adjacency constraints from example inputs to generate larger outputs.

By Dipika Rajesh, Ahmed Khalifa, Julian Togelius
Hugging Face Trending Papers
Jul 6

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms to incorporate constraint information during the generation process.

arXiv AI
Sep 23

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.

By Maxim Gumin, Hsueh-Ti Derek Liu, Victor Zordan, Daniel Ritchie
Hugging Face Trending Papers
Jun 24

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead.

arXiv AI
Jul 7

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

arXiv:2607. 04854v1 Announce Type: new Abstract: Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications.

By Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu