arXiv AI

When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models

arXiv AI
Jun 30

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.

By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv Machine Learning
Sep 22

Characterizing Model-Native Skills

arXiv:2604.17614v2 Announce Type: replace-cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterization...

By Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv AI
3d ago

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.

By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
Hugging Face Trending Papers
Aug 9

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.