Introducing SafeCoder
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2608.30025v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distingui...
arXiv:2606. 11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code.
InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.
arXiv:2610.02300v1 Announce Type: new Abstract: Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, app...
SafeRI proposes an on-demand safety alignment approach for large vision-language models, contrasting with existing always-on methods that globally modify model behavior. The framework uses a lightweight recognizer to evaluate token-level safety during autoregressive generation, gating a LoRA module that only activates when unsafe content is detected. By training the LoRA on unsafe prefixes and safe continuations, SafeRI redirects unsafe generations back to safe responses without perturbing the model’s original reasoning path.