arXiv:2609.36544v1 Announce Type: cross
Abstract: Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through whic...
By Divyansh Chandarana, Sandipan De, Vivek Gupta
High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident.
Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.
By Michal \v{S}tef\'anik, Jan Nehyba, Jirina Karasova, Martin Fico, Lucie \v{S}karkov\'a, Mark\'eta Ko\v{s}atkov\'a, David Kosatka
The article examines how three elementary teachers implemented a conversational AI‑based curriculum using the ToyTalk platform during a three‑week summer camp. Over 13 instructional days, teachers employed adaptive practices—repair, differentiation, translation, and balancing—to navigate tensions among technology, learners, and instruction. Their understanding of AI and instructional roles evolved throughout the camp, leading to design implications for deploying conversational AI in elementary classrooms.
By Fasika Melese, Ruiyang Wu, Xinyue Cui, Joanna Perkins, Xiaoyi Tian, Tiffany Barnes, Shiyan Jiang
arXiv:2606. 00038v1 Announce Type: cross Abstract: Artificial intelligence (AI) literacy is increasingly recognized as a foundational competency for all university graduates.
By J. Paul Liu, Rachel Levy
The study explores how undergraduate computing students in Saudi Arabia perceive AI‑generated writing feedback when they are explicitly told that ChatGPT, not a human instructor, produced the score and comments. Through qualitative reflections, four themes emerged: students found the feedback useful for surface‑level revisions, recognized AI’s contextual and pedagogical limits, trusted the feedback conditionally—separating its utility from its authority—and reaffirmed the human instructor’s role as the ultimate grading authority. The findings highlight a clear distinction students make between feedback usefulness and evaluative authority, treating them as separate judgments rather than opposing ends of a single approval scale.
By Rayed AlGhamdi