arXiv AI

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

arXiv:2605. 17110v2 Announce Type: replace Abstract: Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation.

Hugging Face Trending Papers
Aug 9

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.

arXiv AI
Jun 30

SemJoin: Semantic Join Optimization

arXiv:2606. 29532v1 Announce Type: cross Abstract: Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis.

By Christopher Gou, Aditya Banerjee, Jiaxuan Wang, Chunwei Liu
arXiv AI
Sep 12

Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

The paper introduces Debate‑to‑Skill, a capability‑bound process supervision method for annotating industrial query‑to‑agent tasks. It addresses failures caused by confusing topical relevance with executable capability, especially for long‑tail and boundary‑sensitive requests. By employing reusable decision principles, structured deliberation, verifier‑based verdict extraction, and disagreement‑driven refinement, the approach is evaluated against direct‑label supervision, reasoning‑SFT, and structural ablations on a Query2Agent benchmark, focusing on grey‑zone cases where semantic relatedness and executable capability diverge.

By Shiyu Zhang, Leisheng Cheng, Huifu Li