arXiv AI By Shiyu Zhang, Leisheng Cheng, Huifu Li

Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

Read the original on arXiv AI →

The paper introduces Debate‑to‑Skill, a capability‑bound process supervision method for annotating industrial query‑to‑agent tasks. It addresses failures caused by confusing topical relevance with executable capability, especially for long‑tail and boundary‑sensitive requests. By employing reusable decision principles, structured deliberation, verifier‑based verdict extraction, and disagreement‑driven refinement, the approach is evaluated against direct‑label supervision, reasoning‑SFT, and structural ablations on a Query2Agent benchmark, focusing on grey‑zone cases where semantic relatedness and executable capability diverge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 9

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.

arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller