arXiv:2606. 01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback.
By Aitor Arronte Alvarez, Naiyi Xie Fincham
arXiv:2605. 05598v2 Announce Type: replace Abstract: The proliferation of large language models (LLMs) in educational settings has paradoxically undermined the cognitive processes they purport to support.
By Ran Bi, Shiyao Wei, Yuanyiyi Zhou
The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.
By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
The paper investigates how well large language models (LLMs) can handle character attacks—ad hominem arguments—in political debates. By analyzing natural political dialogues and comparing LLM-generated responses to a corpus of U.S. presidential debates, the study finds that most LLMs favor logical defenses and rarely use ethos-based counterattacks. The authors suggest that safety fine‑tuning limits LLMs’ strategic options, preventing them from fully engaging in realistic political discourse.
By Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak
arXiv:2609.24369v1 Announce Type: cross
Abstract: Deception plays a central role in Intelligence operations, yet it remains difficult to analyse systematically without expert knowledge of reasoning p...
By Stefan Sarkadi, Xabier Garmendia, Jack Mumford, Trevor Bench-Capon
arXiv:2608.22993v1 Announce Type: new
Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...
By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
arXiv:2608. 07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers.
By Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky, Tanja K\"aser
arXiv:2606. 12767v1 Announce Type: new Abstract: Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use.
By Sarah Elshabrawy, Rahul K. Dass, Ashok K. Goel
SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.
By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
ARGUS is a new agent-based framework for persuasive argument generation that incorporates a Theory-of-Mind Reasoner to model audience beliefs and values. It uses a component-aware planner to break arguments into subtopics, assign rhetorical functions (logos, pathos, ethos, kairos), and guide evidence retrieval during planning. A refinement module iteratively addresses multi-dimensional weaknesses, and evaluations on three benchmarks show ARGUS outperforming strong baselines and effectively shifting resistant audience stances.
By Zhe Hu
arXiv:2608. 11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today.
By Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski
arXiv:2602. 02414v2 Announce Type: replace-cross Abstract: Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors.
By Joshua Mitton, Prarthana Bhattacharyya, Digory Smith, Thomas Christie, Ralph Abboud, Simon Woodhead