arXiv Machine Learning By Jiayi Li, Kun Zhan

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs

Read the original on arXiv Machine Learning →

arXiv:2607. 09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 22

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while moderation policies may vary across products, regions, and deployment stages.

arXiv AI
Sep 18

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

The paper investigates whether large language models (LLMs) can internally detect harmful content, bypassing external guardrails that add latency and computational cost. By extracting activations from LLaMA‑3.1‑8B and training lightweight MLP probes, the authors achieve high F1 scores (99%, 83%, and 84%) on WildJailbreak, Beavertails, and AEGIS 2.0 benchmarks, rivaling much larger guard models while reducing overhead. This suggests that internal state monitoring can provide efficient safety checks for resource‑constrained, time‑critical deployments.

By Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi