arXiv AI

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

arXiv AI
Aug 5

AI Assistance Reduces Persistence and Hurts Independent Performance

arXiv:2604. 04721v3 Announce Type: replace Abstract: People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results.

By Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey
arXiv Machine Learning
Sep 21

Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

The paper introduces a new reward, GRPO, that encourages large reasoning models (LRMs) to efficiently determine whether a task is solvable before generating a full chain of thought. Fine‑tuning 4B LRMs with this reward improves their ability to abstain from answering unanswerable prompts by an average of 12.8% while producing 44% shorter chains of thought. The approach also preserves the models’ overall answering performance.

By Polina Tsvilodub, Max H\"oth, Michael Franke, Bj\"orn Deiseroth, Carina Kauf
arXiv AI
Jul 24

AI Assistants Overassist

arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.

By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv AI
Aug 20

What is Missing from AI Post-Training AI: An Empirical Analysis

The paper investigates the limitations of post-training AI agents that can autonomously train large language models. It distinguishes between execution-level capability—making adjustments within a chosen training strategy—and strategy-level capability—revising the overall approach based on new evidence. Analysis of many public post-training runs shows that agents lock into a strategy early and then only perform local tweaks, regardless of task. Experiments with experience scaffolds, human guidance, and extra compute improve execution but do not enable strategy reevaluation, indicating that agents lack a mechanism to spontaneously reassess their strategy during training.

By Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin
arXiv AI
Aug 28

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.

By Md Jueal Mia, M. Hadi Amini
arXiv Machine Learning
Sep 24

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

The paper investigates how large language models (LLMs) balance capability and efficiency when using Chain-of-Thought reasoning on arithmetic and algorithmic tasks. It finds that while larger models solve more problems correctly, the improvement follows an exponential decay that slows with scale, indicating diminishing returns. Additionally, the length of reasoning output grows with problem size but does not improve with larger models, suggesting efficiency does not benefit from scaling.

By Moritz Laber, Zohair Shafi, Germans Savcisens, Brennan Klein, Matteo Chinazzi, Samuel V. Scarpino, Albert-L\'aszl\'o Barab\'asi, Tina Eliassi-Rad