Hugging Face Trending Papers

Self-Evolving Deep Research via Joint Generation and Evaluation

Read the original on Hugging Face Trending Papers →

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jun 19

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments

arXiv:2606. 19893v1 Announce Type: new Abstract: Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning.

By Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li
arXiv Computation and Language
Sep 3

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

The paper introduces a method for creating query‑specific rubrics for DeepResearch‑style long‑form report generation by training rubric generators with reinforcement learning. It builds a dataset of queries annotated with human preferences, then uses a hybrid reward that includes preference consistency, format validity, and LLM‑based rubric evaluation. The learned rubrics outperform generic or manually constructed alternatives in distinguishing preferred reports and, when used as rewards, improve performance of both single‑agent and multi‑agent DeepResearch systems.

By Changze Lv, Jie Zhou, Wentao Zhao, Jingwen Xu, Shihan Dou, Zisu Huang, Muzhao Tian, Xiaohua Wang, Zhengkang Guo, Yang Liu, Pluto Zhou, Tao Gui, Le Tian, Xiao Zhou, Xiaoqing Zheng, Xuanjing Huang, Jie Zhou