arXiv AI By M P V S Gopinadh

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

Read the original on arXiv AI →

This study investigates whether large language models (LLMs) exhibit different safety vulnerabilities when prompted with emojis instead of plain text. By testing 50 emoji‑augmented prompts on four open‑source LLMs—Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B—the authors found varying success rates of unsafe responses: Gemma 2 9B and Mistral 7B each had a 10% success rate, Llama 3 8B had 6%, while Qwen 2 7B was fully resistant. A chi‑square test confirmed significant differences in the outcome distributions, suggesting that input representation can markedly affect model robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Evaluating Language Model Safety Across Long Adversarial Conversations

The paper investigates how conversational safety in language models degrades over extended, adversarial interactions. By testing three instruction‑tuned models with persistent adversarial users across up to 101 turns, the study finds that safe‑response rates drop sharply from 85–100% at the first turn to 15–44% by the end. This demonstrates that strong single‑turn safety does not guarantee continued safety in long conversations.

By Parisa Salmani, Peter R. Lewis
arXiv AI
Sep 4

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark includes 7,200 adversarial prompts covering ten safety-critical content categories, six persuasive strategies, and four languages (Hindi, Bengali, Marathi, Punjabi). Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, revealing that current English-centric safety evaluations miss important multilingual vulnerabilities.

By Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
Hugging Face Trending Papers
Sep 3

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark covers ten safety-critical content categories, six persuasive strategies, and four languages—Hindi, Bengali, Marathi, and Punjabi—producing 7,200 adversarial prompts. Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, highlighting gaps in current English-centric safety assessments.

arXiv AI
4d ago

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r