HintMiner is an automatic tool that mines question hints from web Q&A posts using a language‑model‑based MiningNet. It retrieves many Q&A posts, extracts hints via a transformer‑based encoder‑decoder with copying mechanisms, and is trained with a self‑supervised objective on large online data. Evaluated on 60,000 Stack Overflow questions, HintMiner achieves an average BLEU score of 36.17% and ROUGE‑2 of 36.29%.
By Zhenyu Zhang, JiuDong Yang
Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.
By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
arXiv:2504. 07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness.
By Sher Badshah, Ali Emami, Hassan Sajjad
arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.
By Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King
arXiv:2606. 27047v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge.
By Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims, Selma Peterson, Emily Taylor
arXiv:2606. 06197v1 Announce Type: cross Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs).
By Hafez Abdelghaffar, Ahmed Alansary, Ali Hamdi
arXiv:2607.00711v2 Announce Type: cross
Abstract: Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input require...
By Zheng Fang, Dongming Jin, Yihong dong, Yongmin Li, Kechi Zhang, Zhi Jin, Ge Li
arXiv:2411.10163v3 Announce Type: replace
Abstract: Large language models (LLMs) demonstrate remarkable performance across various tasks, prompting researchers to develop diverse evaluation benchmark...
By Yutao Hou, Yajing Luo, Zhiwen Ruan, Hongru Wang, Weifeng Ge, Yun Chen, Guanhua Chen
The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.
By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach
arXiv:2605. 26937v2 Announce Type: replace-cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood.
By Luca Giordano, Simon Razniewski
arXiv:2604. 09497v2 Announce Type: replace-cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.
By Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Emmanuel Malherbe, C\'eline Hudelot, Pierre Colombo
CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.
By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li