arXiv:2607. 28008v2 Announce Type: replace-cross Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data.
By Yanshi Li, Xueru Bai, Shuman Liu, Long Zhang
arXiv:2606. 26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run.
By Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Ant\'ia Garc\'ia, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay
arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
By Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
By Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.
arXiv:2606. 29532v1 Announce Type: cross Abstract: Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis.
By Christopher Gou, Aditya Banerjee, Jiaxuan Wang, Chunwei Liu
arXiv:2606. 19079v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters.
By Enrico Cassano, Micha{\l} Brzozowski, Paolo Mandica, Zuzanna Dubanowska, Neo Christopher Chung
The paper introduces Debate‑to‑Skill, a capability‑bound process supervision method for annotating industrial query‑to‑agent tasks. It addresses failures caused by confusing topical relevance with executable capability, especially for long‑tail and boundary‑sensitive requests. By employing reusable decision principles, structured deliberation, verifier‑based verdict extraction, and disagreement‑driven refinement, the approach is evaluated against direct‑label supervision, reasoning‑SFT, and structural ablations on a Query2Agent benchmark, focusing on grey‑zone cases where semantic relatedness and executable capability diverge.
By Shiyu Zhang, Leisheng Cheng, Huifu Li
arXiv:2606. 12451v1 Announce Type: new Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck.
By Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber, Sahil Bansal
arXiv:2605.05726v2 Announce Type: replace
Abstract: As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical...
By Hongcheol Cho, Ryangkyung Kang, Youngeun Kim
arXiv:2606. 06924v1 Announce Type: new Abstract: Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers.
By Guannan Lai, Haoran Hu, Long Chen, Zhenguo Li, Han-Jia Ye
arXiv:2602.13540v2 Announce Type: replace-cross
Abstract: Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on...
By Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun