arXiv:2606. 01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback.
By Aitor Arronte Alvarez, Naiyi Xie Fincham
arXiv:2605. 05598v2 Announce Type: replace Abstract: The proliferation of large language models (LLMs) in educational settings has paradoxically undermined the cognitive processes they purport to support.
By Ran Bi, Shiyao Wei, Yuanyiyi Zhou
arXiv:2608. 07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers.
By Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky, Tanja K\"aser
arXiv:2606. 12767v1 Announce Type: new Abstract: Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use.
By Sarah Elshabrawy, Rahul K. Dass, Ashok K. Goel
arXiv:2608. 11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today.
By Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski
arXiv:2602. 02414v2 Announce Type: replace-cross Abstract: Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors.
By Joshua Mitton, Prarthana Bhattacharyya, Digory Smith, Thomas Christie, Ralph Abboud, Simon Woodhead
arXiv:2606. 15766v1 Announce Type: new Abstract: A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution.
By Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson
arXiv:2606. 26698v1 Announce Type: cross Abstract: In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth of information disorder.
By Eleni Papadopulos, Firoj Alam, Giovanni Da San Martino
Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from decoding, can predict reasoning quality as assessed by LLM-as-judge evaluation.
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv:2606. 12754v1 Announce Type: cross Abstract: Are large language models (LLMs) bad at capturing human judgment?
By Danica Dillion, Chen Cecilia Liu, Baihui Wang, Daniele Barolo, Tanmay Rajore, Niket Tandon, Pranathi Ravikumar, Kurt Gray
arXiv:2608. 17919v1 Announce Type: cross Abstract: Background and Context: Question and inquiry are integral parts of knowledge seeking and learning.
By Matin Amoozadeh, Amin Alipour