arXiv AI By Mohammed Yousif, Prabhjot Singh, Arjun Pankajakshan, Madhu Reddiboina

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

Read the original on arXiv AI →

arXiv:2607. 22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization

PaperScout is an autonomous agent that treats academic paper search as a sequential decision-making process, dynamically deciding when and how to use search and expansion tools based on accumulated context. The authors identify a granularity mismatch in standard reinforcement learning for multi-turn tasks and propose Proximal Sequence Policy Optimization (PSPO), a sequence-level policy optimization method that aligns learning with agent–environment interactions. Experiments on synthetic and real-world benchmarks show that PaperScout outperforms structured retrieval and RL baselines in recall and relevance, demonstrating the effectiveness of its adaptive agentic framework and optimization strategy.

By Tingyue Pan, Jie Ouyang, Mingyue Cheng, Qingchuan Li, Zirui Liu, Daoyu Wang, Mingfan Pan, Shuo Yu, Qi Liu, Enhong Chen