arXiv AI

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

arXiv:2607. 22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type.

arXiv AI
Sep 4

PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization

PaperScout is an autonomous agent that treats academic paper search as a sequential decision-making process, dynamically deciding when and how to use search and expansion tools based on accumulated context. The authors identify a granularity mismatch in standard reinforcement learning for multi-turn tasks and propose Proximal Sequence Policy Optimization (PSPO), a sequence-level policy optimization method that aligns learning with agent–environment interactions. Experiments on synthetic and real-world benchmarks show that PaperScout outperforms structured retrieval and RL baselines in recall and relevance, demonstrating the effectiveness of its adaptive agentic framework and optimization strategy.

By Tingyue Pan, Jie Ouyang, Mingyue Cheng, Qingchuan Li, Zirui Liu, Daoyu Wang, Mingfan Pan, Shuo Yu, Qi Liu, Enhong Chen
arXiv AI
3d ago

GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.

By Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i}