← Back to all news
arXiv Machine Learning August 31, 2026 By Szymon Mi{\l}osz, Piotr Duch, Szymon Grabowski

Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning
  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 5

When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero

arXiv:2504. 14636v3 Announce Type: replace-cross Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search.

By Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
agentsreinforcement-learningefficiencybenchmarks
More like this →
arXiv AI
Jul 9

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

arXiv:2607. 06854v1 Announce Type: cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade, since they beat a random opponent over 99 percent of the time and only tie copies of themselves.

By Nima Kelidari, Mohammadsaeed Haghi, Mahdi Salmani
llmsragagentsreinforcement-learning
More like this →
arXiv AI
Jun 11

The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network

arXiv:2508. 21380v3 Announce Type: replace-cross Abstract: Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents.

By Elias Sandmann, Sebastian Lapuschkin, Wojciech Samek
agentssafety
More like this →
arXiv AI
Aug 14

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

arXiv:2608. 12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment.

By Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li
agentsreinforcement-learningefficiency
More like this →
arXiv AI
3d ago

CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

arXiv:2609.00892v1 Announce Type: new Abstract: Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforceme...

By Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu, Ke Zeng, Jinli Suo
llmsreinforcement-learningbenchmarks
More like this →
arXiv AI
Aug 18

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

arXiv:2608. 14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice.

By Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, Chris Piech
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea