arXiv AI

PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms

arXiv:2603. 27476v2 Announce Type: replace Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance.

arXiv AI
Sep 1

PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded Verification

arXiv:2603.27476v3 Announce Type: replace Abstract: AI-powered people search platforms are increasingly deployed for recruiting, sales prospecting, and professional networking, yet no standardized be...

By Tianyu Shi, Wei Wang, Zequn Xie, Shuai Zhang, Boyang Xia, Chenyu Zeng, Qi Zhang, Lynn Ai, Yaqi Yu, Kaiming Zhang, Feiyue Tang, Zhenyu Yu, Lei Ding
arXiv AI
Jun 12

DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks

arXiv:2606. 12871v1 Announce Type: new Abstract: Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses.

By Jingxuan Han, Wei Liu, Mingyang Zhu, Youpeng Wang, Ziwen Wang, Lin Qiu, Xuezhi Cao, Xunliang Cai, Zheren Fu, Licheng Zhang, Zhendong Mao
arXiv AI
Aug 24

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.

By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
Hugging Face Trending Papers
Sep 8

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.

arXiv AI
Sep 25

SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure‑mode attribution. It employs a two‑stage training pipeline to align evaluation with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on industrial short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and significantly advances attribution diagnosis, leading to its deployment at Kuaishou with over 400 million daily active users.

By Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo
arXiv AI
Sep 12

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Benchmark Radar is a living database and search engine that aggregates AI benchmark papers, datasets, code, and score histories. It automatically discovers new benchmark resources from 37 sources, maintains a catalog of 1,283 records with 12,916 numeric observations, and provides tools such as a web dashboard, CLI, and downloadable evidence for researchers. The system also offers visualizations like a Pareto frontier and trend views to help users assess benchmark saturation and adoption.

By Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu
arXiv AI
Sep 18

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

The study examines how conversational LLM agents—specifically ChatGPT, Claude, Grok, and DeepSeek—use Web search, combining real user interactions with controlled API experiments. It finds that agents differ in when they decide to search, how they craft queries, and which domains they favor, and that more frequent searching does not always improve answer quality. While most responses are grounded in search results, some claims are unsupported, raising attribution concerns.

By Mahsa Amani, Seungeon Lee, Abhisek Dash, Asmaa El Fraihi, Yunah Jang, Elisabeth Kirsten, Qinyuan Wu, Krishna P. Gummadi, Manish Gupta, Abhilasha Ravichander, Muhammad Bilal Zafar, Soumi Das
Hugging Face Trending Papers
Sep 24

SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure mode attribution. It employs a two‑stage training pipeline to align the evaluator with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on Kuaishou’s short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and enhances attribution diagnosis, leading to better online search evaluation at scale.