arXiv Computation and Language

HintEval: An Open-Source Python Toolkit for Hint Generation and Hint Evaluation

arXiv AI
Sep 16

HintMiner: Automatic Question Hints Mining From Q&A Web Posts with Language Model via Self-Supervised Learning

HintMiner is an automatic tool that mines question hints from web Q&A posts using a language‑model‑based MiningNet. It retrieves many Q&A posts, extracts hints via a transformer‑based encoder‑decoder with copying mechanisms, and is trained with a self‑supervised objective on large online data. Evaluated on 60,000 Stack Overflow questions, HintMiner achieves an average BLEU score of 36.17% and ROUGE‑2 of 36.29%.

By Zhenyu Zhang, JiuDong Yang
arXiv AI
6d ago

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.

By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
arXiv AI
Jul 21

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.

By Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King
arXiv Machine Learning
Sep 21

Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.

By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach
arXiv AI
2d ago

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.

By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li