Hugging Face Trending Papers

SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure mode attribution. It employs a two‑stage training pipeline to align the evaluator with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on Kuaishou’s short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and enhances attribution diagnosis, leading to better online search evaluation at scale.

arXiv AI
Sep 25

SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure‑mode attribution. It employs a two‑stage training pipeline to align evaluation with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on industrial short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and significantly advances attribution diagnosis, leading to its deployment at Kuaishou with over 400 million daily active users.

By Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo
arXiv AI
Sep 1

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.

By Sahil Kale
arXiv AI
Aug 7

Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning

arXiv:2608. 05245v1 Announce Type: new Abstract: Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains.

By Muyang Ye, Tian Lan, Feihu Jiang, Yongshi Ye, Wuyunsiqin, Bin Zhu, Qianghuai Jia, Zhao Xu, Weihua Luo, Ye Wang, Jinyang Zhang, Longyue Wang, Lingfeng Bao
arXiv Computation and Language
Sep 17

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

M‑SQE is a post‑retrieval framework that estimates the quality of multilingual agent skills by combining a Theory view (intrinsic quality) and an Action view (task‑grounded utility) into a domain‑conditioned score. It was evaluated on general, tool‑use, and cultural skill‑use domains, showing a task‑success improvement of at least +3.5 points over baselines across three retrievers. The method notably boosts performance for low‑resource languages, raising Hindi by +12.9 pp and Swahili by +5.6 pp, and achieves strong results across six cultural regions, advancing linguistic and cultural equality in agentic skill use.

By Yilun Liu, Shimin Tao, Minggui He, Chenxin Liu, Li Zhang, Chen Liu, Miao Zhang, Jiaxin Guo, Min Zhang, Liqun Deng, Xiaojun Meng, Daimeng Wei
arXiv AI
Jun 17

A Framework for Evaluating Agentic Skills at Scale

arXiv:2606. 17819v1 Announce Type: cross Abstract: Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and no reusable methodology exists for evaluating an individual skill.

By Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich, Maria I. Gorinova, Amy Heineike, Rob Willoughby
arXiv AI
Sep 3

PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation

PRO-Step introduces a step‑level process reward optimization framework for Retrieval‑Augmented Generation (RAG) that evaluates both logical validity and evidential grounding at each reasoning step. By training a generative Preference‑Based Reward Model (PRM) and using PRM‑guided value tree search to create preference pairs, the method optimizes the policy through step‑level Direct Preference Optimization. Experiments on single and multi‑hop QA benchmarks show that PRO‑STEP achieves the best average EM and F1 scores across five datasets.

By MinKeon Kim, Namjun Lee, Jaekwang Kim