SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure‑mode attribution. It employs a two‑stage training pipeline to align evaluation with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on industrial short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and significantly advances attribution diagnosis, leading to its deployment at Kuaishou with over 400 million daily active users.
By Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo
arXiv:2605. 09038v3 Announce Type: replace Abstract: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries.
By Jinchao Hu, Meizhi Zhong, Kehai Chen, Min Zhang
arXiv:2608. 09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge.
By Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang
The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.
By Sahil Kale
arXiv:2608. 05245v1 Announce Type: new Abstract: Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains.
By Muyang Ye, Tian Lan, Feihu Jiang, Yongshi Ye, Wuyunsiqin, Bin Zhu, Qianghuai Jia, Zhao Xu, Weihua Luo, Ye Wang, Jinyang Zhang, Longyue Wang, Lingfeng Bao
M‑SQE is a post‑retrieval framework that estimates the quality of multilingual agent skills by combining a Theory view (intrinsic quality) and an Action view (task‑grounded utility) into a domain‑conditioned score. It was evaluated on general, tool‑use, and cultural skill‑use domains, showing a task‑success improvement of at least +3.5 points over baselines across three retrievers. The method notably boosts performance for low‑resource languages, raising Hindi by +12.9 pp and Swahili by +5.6 pp, and achieves strong results across six cultural regions, advancing linguistic and cultural equality in agentic skill use.
By Yilun Liu, Shimin Tao, Minggui He, Chenxin Liu, Li Zhang, Chen Liu, Miao Zhang, Jiaxin Guo, Min Zhang, Liqun Deng, Xiaojun Meng, Daimeng Wei