arXiv AI

Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical Argumentation

arXiv:2606. 01020v1 Announce Type: new Abstract: Identifying logical fallacies in everyday discourse is challenging for many people.

arXiv Computation and Language
Sep 1

Evaluating the Capabilities of LLMs for Persuasive Dialogue

The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.

By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
arXiv Computation and Language
Sep 25

Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

The paper investigates how well large language models (LLMs) can handle character attacks—ad hominem arguments—in political debates. By analyzing natural political dialogues and comparing LLM-generated responses to a corpus of U.S. presidential debates, the study finds that most LLMs favor logical defenses and rarely use ethos-based counterattacks. The authors suggest that safety fine‑tuning limits LLMs’ strategic options, preventing them from fully engaging in realistic political discourse.

By Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak
arXiv Computation and Language
Aug 25

LLM Pedagogical Behavior in AI Tutoring Interactions

arXiv:2608.22993v1 Announce Type: new Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...

By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
arXiv Computation and Language
Sep 24

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.

By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
arXiv Computation and Language
Aug 24

ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding

ARGUS is a new agent-based framework for persuasive argument generation that incorporates a Theory-of-Mind Reasoner to model audience beliefs and values. It uses a component-aware planner to break arguments into subtopics, assign rhetorical functions (logos, pathos, ethos, kairos), and guide evidence retrieval during planning. A refinement module iteratively addresses multi-dimensional weaknesses, and evaluations on three benchmarks show ARGUS outperforming strong baselines and effectively shifting resistant audience stances.

By Zhe Hu