Training Documents Reranker with Search Rubrics for Deep Research Agent
arXiv:2608. 03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents.
The paper introduces a method for creating query‑specific rubrics for DeepResearch‑style long‑form report generation by training rubric generators with reinforcement learning. It builds a dataset of queries annotated with human preferences, then uses a hybrid reward that includes preference consistency, format validity, and LLM‑based rubric evaluation. The learned rubrics outperform generic or manually constructed alternatives in distinguishing preferred reports and, when used as rewards, improve performance of both single‑agent and multi‑agent DeepResearch systems.
arXiv:2608. 03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents.
Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning.
The paper introduces F$^{2}$DR, a fine‑grained reward framework designed to evaluate end‑to‑end DeepSearch workflows, which involve planning, reflection, retrieval, and answer generation. F$^{2}$DR assesses workflows along three dimensions—Content, Trajectory, and Answer—to provide a comprehensive process‑level evaluation. The authors also present DeepSearch RM‑Bench, a benchmark that tests reward models in DeepSearch scenarios and shows strong discriminative power over existing open‑source models.
arXiv:2606. 04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability.
arXiv:2610.00389v1 Announce Type: cross Abstract: Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited infor...
arXiv:2607. 24850v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons.
Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools. However, existing agents often incur substantial computational and interaction costs, generating lengthy trajectories that contain redundant queries, inefficient exploration, and irrelevant observations.
arXiv:2605. 01248v3 Announce Type: replace Abstract: Reinforcement learning (RL) post-training has enabled newer capabilities in models, such as agentic tool-use for search.
arXiv:2606. 27291v1 Announce Type: new Abstract: Job-search platforms rely on low-bandwidth query interfaces that often fail to capture the high-dimensional complexity of candidate profiles.
arXiv:2608.29856v1 Announce Type: new Abstract: Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during...
arXiv:2606. 12871v1 Announce Type: new Abstract: Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses.
arXiv:2601. 07055v2 Announce Type: replace Abstract: As high-quality data becomes increasingly difficult to obtain, self-evolution without curated training data has emerged as a promising paradigm.