Introducing SafeCoder
Related stories
Interpreting and Steering for Safe and Correct Code Generation
arXiv:2608.30025v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distingui...
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
arXiv:2606. 11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code.
InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation
InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.
Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
arXiv:2610.02300v1 Announce Type: new Abstract: Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, app...
SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models
SafeRI proposes an on-demand safety alignment approach for large vision-language models, contrasting with existing always-on methods that globally modify model behavior. The framework uses a lightweight recognizer to evaluate token-level safety during autoregressive generation, gating a LoRA module that only activates when unsafe content is detected. By training the LoRA on unsafe prefixes and safe continuations, SafeRI redirects unsafe generations back to safe responses without perturbing the model’s original reasoning path.
Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions
arXiv:2607. 08011v1 Announce Type: cross Abstract: Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code.
Introducing the LiveCodeBench Leaderboard - Holistic and Contamination-Free Evaluation of Code LLMs
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
The paper introduces Safety-aware Contrastive Decoding (SafeCoDe), a lightweight, model‑agnostic framework designed to improve context‑aware safety in Multimodal Large Language Models (MLLMs). SafeCoDe operates in two stages: a contrastive decoding step that highlights tokens sensitive to visual context by contrasting real and Gaussian‑noised images, and a global‑aware token modulation strategy that adjusts refusals based on scene‑level reasoning and predicted safety verdicts. Experiments across various MLLM architectures and safety benchmarks demonstrate that SafeCoDe consistently enhances context‑sensitive refusal behaviors while maintaining model helpfulness.
ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
arXiv:2609.39688v1 Announce Type: cross Abstract: Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment...
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
arXiv:2502. 02068v3 Announce Type: replace-cross Abstract: This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development.
Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
arXiv:2606. 05290v1 Announce Type: cross Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture.