An Introduction to AI Secure LLM Safety Leaderboard
Related stories
Bringing the Artificial Analysis LLM Performance Leaderboard to Hugging Face
Open LLM Leaderboard: DROP deep dive
Online Safety Monitoring for LLMs
arXiv:2607. 02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time.
Introducing the Open FinLLM Leaderboard
Democratizing AI Safety with RiskRubric.ai
Introducing the Open Leaderboard for Japanese LLMs!
Fixing Open LLM Leaderboard with Math-Verify
Operator System Card
Drawing from OpenAI’s established safety frameworks, this document highlights our multi-layered approach, including model and product mitigations we’ve implemented to protect against prompt engineering and jailbreaks, protect privacy and security, as well as details our external red teaming efforts, safety evaluations, and ongoing work to further refine these safeguards.
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
arXiv:2606. 20408v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.
Our approach to AI safety
Ensuring that AI systems are built, deployed, and used safely is critical to our mission.