arXiv Machine Learning

WANDR: A Benchmark for Wide and Deep Research

arXiv:2608. 14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents.

arXiv Computation and Language
Sep 11

DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.

By Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao
arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee
arXiv AI
Sep 7

Iris: Climbing to the Search Frontier

The paper introduces Iris-mini and Iris-pro, two search agents trained at 35B and 397B parameter scales. They use a novel data pipeline that constructs reverse‑engineered multi‑hop queries from web hyperlinks, filters trajectories, and alternates supervised fine‑tuning with reinforcement learning in a process called SFT‑RL climbing. Evaluations on several benchmarks show that, with inference‑time context management, the agents achieve the best open‑source results in their parameter ranges.

By Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan
arXiv AI
Sep 18

Efficiently Linking Unstructured Data for Multi-step Reasoning

The paper introduces DASE, a query engine designed to efficiently link unstructured data for multi-step reasoning tasks. DASE combines a multi-step reasoning model, a sparse materialized embedding-similarity join index (SemJI), and a co‑designed execution layer to perform multi‑attribute filtering, multi‑vector search, exact relational joins, and thresholded embedding‑similarity joins. In scientific discovery workloads, DASE outperforms traditional RDBMS, rerank, and vector‑database baselines by 6x to 46x in retrieval speed while maintaining comparable recall, and it serves as a high‑recall prefilter that reduces downstream LLM evaluation cost and improves accuracy on benchmarks such as SemBench E‑Commerce.

By Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar, Zachary Ives
arXiv AI
Aug 24

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.

By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
arXiv Computation and Language
Aug 25

Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

arXiv:2608.22817v1 Announce Type: new Abstract: Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure...

By Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour, Masoud Nasiripour, Moharram Challenger
arXiv AI
Sep 12

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Benchmark Radar is a living database and search engine that aggregates AI benchmark papers, datasets, code, and score histories. It automatically discovers new benchmark resources from 37 sources, maintains a catalog of 1,283 records with 12,916 numeric observations, and provides tools such as a web dashboard, CLI, and downloadable evidence for researchers. The system also offers visualizations like a Pareto frontier and trend views to help users assess benchmark saturation and adoption.

By Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu
arXiv AI
Sep 23

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim as...

By ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin
arXiv AI
Jun 9

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

arXiv:2606. 08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence.

By Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou, Wenjing Gong, Kaili Zhang, Bingqian Chen, Mitch Zhang, Yifan Yang