← Back to all news
Hugging Face Trending Papers September 1, 2026

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • fine-tuning
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Sep 2

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

arXiv:2609.01046v1 Announce Type: cross Abstract: Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigat...

By Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin
llmsfine-tuningbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 3

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

arXiv:2607. 02079v1 Announce Type: cross Abstract: We present HaloGuard 1.

By Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin
agentsbenchmarkssafety
More like this →
arXiv Computation and Language
Aug 25

BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models

arXiv:2608.21880v1 Announce Type: new Abstract: Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely wri...

By Md. Rakibul Hassan, Muhammad Iqbal Hossain
llmsmultimodalbenchmarkssafety
More like this →
arXiv AI
Jul 20

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.

By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
llmsmultimodalbenchmarkssafety
More like this →
arXiv AI
Jul 28

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills

arXiv:2604. 06550v3 Announce Type: replace-cross Abstract: Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access.

By Yinghan Hou, Zongyou Yang
llmsagentsbenchmarkssafety
More like this →
arXiv AI
Aug 11

SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

arXiv:2608. 08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored.

By Xinze Chen, Chi Zhang, Ping Ji, Yimin Liu
llmsagentsroboticssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea