Hugging Face Trending Papers

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

The paper extends a dynamical systems approach to classify unsafe outputs from large language models (LLMs) by projecting prompts and responses into high‑dimensional embeddings and fitting separate Koopman-based predictive models for safe and unsafe regimes. A differential residual score compares prediction errors from these models to classify new outputs. Experiments on three safety benchmarks show that including prompt embeddings improves detection of interaction‑dependent violations, especially with causal decoders like Llama‑3, while response‑only violations benefit more from dense semantic embeddings.

arXiv Machine Learning
Sep 11

Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting

The paper introduces a posterior reweighting framework to explain and counter in-context learning jailbreaks in multimodal large language models. It models the model as switching between safe and harmful behavioral modes, interpreting prompt demonstrations as evidence that shifts the posterior. Using this view, the authors derive scaling laws for jailbreak effectiveness and propose a defense that injects benign counter‑evidence to suppress harmful drift while maintaining utility.

By Xu Zhang, Dev Mistry, Xiang Xu, Ren Wang
arXiv Machine Learning
Jul 30

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.

By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang
arXiv Machine Learning
5d ago

Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction

The paper introduces a low-cost method for detecting hallucinations in large language models by treating the model as a black-box dynamical system. It projects responses into a high-dimensional manifold, models the latent state-space dynamics with Koopman operator theory, and uses differential residual scores from transition operators to distinguish factual from hallucinated outputs. The approach requires only a single-sample pass and shows state-of-the-art performance across three benchmarks with reduced resource overhead.

By Dan Wilson, Mohamed Akrout
arXiv AI
Jun 9

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.

By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.