arXiv AI By Tung-Ling Li, Hongliang Liu, Yuhao Wu

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

Read the original on arXiv AI →

arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.

By Sahil Kadadekar
arXiv AI
Jun 3

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

arXiv:2606. 02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat.

By Alexandre Cristov\~ao Maiorano