Hugging Face Blog
May 4, 2023
Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples
arXiv:2606. 03852v1 Announce Type: cross Abstract: Large language models often generate code with bugs.
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.