arXiv AI

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

arXiv:2608. 14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice.

arXiv AI
3d ago

Conditional Generation of Creative Chess Puzzles with Diffusion Models

The paper presents a method for generating creative chess puzzles using masked diffusion models that can be conditioned on tactical themes and partial board positions. It introduces an auxiliary best‑move prediction task that boosts solution uniqueness by 11.6% and theme‑conditioning accuracy by 2.5%. A reinforcement learning framework further increases the yield of unique, theme‑matching puzzles by 89.1%, and the authors release open‑weights models for the community.

By Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi
arXiv AI
Aug 19

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

The paper presents a tutoring platform that combines a generative AI chatbot with a reinforcement learning algorithm to adaptively sequence practice problems for students learning Python. In a five‑month field study across ten high schools, the adaptive sequencing improved unassisted final exam performance by 0.15 standard deviations, with mediation analysis indicating that higher engagement drove the gains. The study demonstrates that signals from student‑chatbot interactions can be leveraged to personalize and optimize learning at scale.

By Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
arXiv Machine Learning
Sep 23

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.

By Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang