arXiv AI

Training Documents Reranker with Search Rubrics for Deep Research Agent

arXiv:2608. 03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents.

arXiv Computation and Language
Sep 3

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

The paper introduces a method for creating query‑specific rubrics for DeepResearch‑style long‑form report generation by training rubric generators with reinforcement learning. It builds a dataset of queries annotated with human preferences, then uses a hybrid reward that includes preference consistency, format validity, and LLM‑based rubric evaluation. The learned rubrics outperform generic or manually constructed alternatives in distinguishing preferred reports and, when used as rewards, improve performance of both single‑agent and multi‑agent DeepResearch systems.

By Changze Lv, Jie Zhou, Wentao Zhao, Jingwen Xu, Shihan Dou, Zisu Huang, Muzhao Tian, Xiaohua Wang, Zhengkang Guo, Yang Liu, Pluto Zhou, Tao Gui, Le Tian, Xiao Zhou, Xiaoqing Zheng, Xuanjing Huang, Jie Zhou
Hugging Face Trending Papers
Sep 8

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.

arXiv Computation and Language
4d ago

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

AdaTutoRank introduces a setwise document reranker that uses Adaptive Tutoring Optimization (ATO) to provide graded supervision across nine rubric dimensions. By generating hint‑based silver labels, reinforcement rewards, and distillation cues tailored to each rollout’s quality, the method improves credit assignment for individual documents within a set. Experiments on ten benchmarks covering Retrieval‑Augmented Generation (RAG), deep research, and setwise evaluation show that AdaTutoRank achieves superior overall performance while reducing the number of retrieval calls.

By Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi