arXiv:2609.36073v1 Announce Type: cross
Abstract: The rapid proliferation of large language models (LLMs) in the context of education has introduced significant challenges in enforcement of academic...
By David Racovan, Ajay Rawat, Christopher K. May, Jeffrey A. Turkstra
Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios.
arXiv:2508.02312v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine...
By Kang Chen, Xiuze Zhou, Yuanhui Yu, Yuanguo Lin, Hefeng Chen, Congyu Cai, Li Shen
The paper investigates prompt injection attacks on 14 open‑source and 3 closed‑source large language models (LLMs), introducing a new metric called Attack Success Probability (ASP) that accounts for uncertainty in model responses. It demonstrates that a simple hypnotism attack can trigger objectionable behavior in models such as StableLM2, Mistral, Openchat, and Vicuna, achieving roughly 90% ASP. The study highlights that moderately well‑known LLMs are particularly vulnerable, underscoring the importance of public awareness and effective mitigation strategies.
By Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen
arXiv:2607. 10402v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level security challenge.
By Lingwei Wei, Dou Hu, Wei Zhou, Songlin Hu, Philip S. Yu
arXiv:2606. 12737v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources.
By Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le
SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.
By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
arXiv:2606. 31639v1 Announce Type: cross Abstract: Large language models are no longer only text generators.
By Seyed Bagher Hashemi Natanzi, Bo Tang
arXiv:2606. 25059v2 Announce Type: replace-cross Abstract: Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs.
By Lena Libon, Pura Peetathawatchai, Michael Aerni, Daniel Paleka, Florian Tram\`er
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.
arXiv:2607. 18622v1 Announce Type: cross Abstract: Textual Collaborative Prompt Optimization (TCPO) extends Textgrad (Yuksekgonul et al.
By Xinting Liao, Behnoosh Zamanlooy, Masoumeh Shafieinejad, David B. Emerson, Ruinan Jin, Deval Pandya, Xiaoxiao Li
arXiv:2606. 01567v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable skills i.
By Yoshinari Fujinuma, Varun Gangal, Traian Rebedea, Makesh Narasimhan Sreedhar, Prasoon Varshney, Rebecca Qian, Anand Kannappan