arXiv AI By Mohamed Akrout, Olivera Kotevska, Dan Wilson

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

Read the original on arXiv AI →

arXiv:2608. 19579v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 20

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

The paper extends a dynamical systems approach to classify unsafe outputs from large language models (LLMs) by projecting prompts and responses into high‑dimensional embeddings and fitting separate Koopman-based predictive models for safe and unsafe regimes. A differential residual score compares prediction errors from these models to classify new outputs. Experiments on three safety benchmarks show that including prompt embeddings improves detection of interaction‑dependent violations, especially with causal decoders like Llama‑3, while response‑only violations benefit more from dense semantic embeddings.

arXiv Machine Learning
Jul 30

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.

By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang